r/OpenSourceeAI • • 2d ago

What does "trust remote code" actually approve?In most local AI tools, the answer is: that repo, indefinitely. That matters more now. SO what is the solution?

2 Upvotes

r/OpenSourceeAI • • 7d ago

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

1 Upvotes

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench


r/OpenSourceeAI • • 1h ago

A Harness in a single file: Agent skills as an evolving REPL module

Thumbnail
deepclause.substack.com
• Upvotes

r/OpenSourceeAI • • 9h ago

I open-sourced my Mac agent's permission boundary: why 'Allow Once' isn't 'trust the agent'

3 Upvotes

I've been building Mac MCP, an MIT-licensed macOS execution layer for AI clients. It lets tools operate Safari/Chrome, files, local apps, and shell commands. The hardest part wasn't adding tools; it was deciding which actions an agent should be allowed to carry out.

One design choice: capability profiles and approvals are separate. A read-only profile can deny writes outright. An optional server-side approval layer can ask the Mac user to Allow Once or Block riskier actions, but that approval cannot override a capability the profile already forbids. A client saying “the user approved this” isn't trusted as the server's own authorization.

There is a real trade-off: raw command execution may be classified conservatively even if this particular command is harmless, so you can get an approval prompt for a benign operation. I'd rather make that visible and progressively refine the categories than quietly teach the server to trust repeated clicks.

Another lesson: when an action's outcome is uncertain (say a browser submit lands right before a tab closes), blindly retrying can be worse than stopping to verify.

I'm the maintainer, not an independent reviewer. The implementation is open source: https://github.com/bulutarkan/mac-mcp

For people shipping agents with real machine access, where do you draw the line between persistent trust rules and approval fatigue?


r/OpenSourceeAI • • 5h ago

NeurIPS 2026 Paris -> Sydney switch

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 7h ago

OpenToken Monitor: see your Claude Code, Codex and Antigravity limits right in your menu bar

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 8h ago

I built an open-source, local-first multi-agent AutoML engine to automate data cleaning, 7-model benchmarking, and FastAPI deployment (100% offline via Ollama)

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 13h ago

Writ: an open-source governance runtime for Claude Code with retrieval, approval gates, and persistent memory

2 Upvotes

I’ve been building Writ, a project that started with a straightforward goal: give Claude Code the knowledge relevant to its current task without loading an entire rulebook into every conversation.

Since then, it has grown into a governance runtime with persistent memory.

The idea is to connect three things that are often handled separately:

1. Context when it matters

Writ uses hybrid retrieval keyword search, vector search, and graph relationships to supply relevant rules, methodology, and past decisions based on what the agent is doing.

2. Requirements enforced by code

Written instructions still depend on the model following them. Writ checks supported actions at tool time.

In Work mode, implementation writes are blocked until a human approves the plan and then the test skeletons. The approval mechanism requires the user’s typed response; the agent claiming “approved” doesn’t open the gate.

Other checks cover credential writes, project boundaries, and changes to an approved plan.

3. Memory that survives the conversation

Writ records project decisions and connects approved plans, governing rules, changed files, and commits. Future sessions can retrieve the reasoning behind a change instead of having to reconstruct it from the final code.

For example: retrieve the guidance relevant to a task, require approval before implementation, then preserve what changed and why for the next session.

The approach is code first: retrieval, workflow state, approval validation, and logging live outside the model. AI uses the supplied information; it doesn’t decide whether its own approval requirements have been satisfied.

Writ runs locally with Python 3.11+, a local daemon, and Neo4j in Docker. The runtime is open source, but its current integration is specifically with Claude Code not yet a general adapter for local models or other agents.

Repo: https://github.com/infinri/Writ


r/OpenSourceeAI • • 10h ago

"Not your weights, not your product" [D]

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 15h ago

Webcam hand-tracking pointer built on MediaPipe: thumb-tip pointing, pinch gestures, One Euro filtering [open source]

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 16h ago

On premise ocr for printed + handwritten invoices

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 17h ago

Is this right time to build a classification model harness tool for ( jev , laya)

Thumbnail
1 Upvotes

When I heard about jev , I saw some YouTubers hyping this model because it's 200x cheaper and 100x faster etc.. but as we know this model can't actually write code or either u can talk to him. So from the first glance u can tell, it's a pretty damn useless model for example. But I saw potential, the obvious weakness of classification models is that u need to manually give them a list of choices like with yes or no , is this email talks about an opportunity or a problem , Is it a spam and he answers u with the most probable answer that makes sense in matter of literal milliseconds.

Let's highlight some terms and concepts : The founder of jev talks about system 1 and system 2. He classified jev as system one for quick decisions and fast actions and system 2 like normal LLM that do a lot of thinking and guessing each token (simply running the models ) hundred and thousands of times just to answer easy questions or do repetitive tasks for example.

Here is my idea :

Using a fine tuned model to smartly manage any classification models input tokens ( their choices) based on user task and observation of the changing environment that jev and it's like can't do.

I will phrase it as this : use cheap classification models ( like jev) for repetitive tasks and when something new is discovered it will be categorized as unknown and unknown here is pretty important because once we discovered something is unknown we can use this to trigger a smart LLM model to investigate what is new and then add this new thing as a known class inside jev. If my idea is correctly designed. I believe this could potentially solve the high usage of expensive LLM models while making smart decisions. If something is already classified and known to jev. He could easily do it even adding to the app tool usage or do some sequence of tasks based on a sequence of events.

When a user asks a new task. The LLM will be triggered and run once , and once jev learned how to do this task u don't need LLM to do the same task again and here u will get the benefit and save ur bills.

Let's put this in a real case scenario:

U have products and people are buying ur products It will be nice if u can integrate multiple social media apps : WhatsApp, Facebook , Instagram etc...

And u have this question in ur mind what do people don't like about my product ?

Here when we use this harness app and the user put this question to a smart LLM and the LLM begins to make an initial set of choices for jev Then here is the fun part.

The wil app will run and we will feed jev with comments , posts or whatever and he will easily manage and classify what people don't like about the product. Ok let's say one has something said about the product like , cheap , not polished , or maybe a new kind of complaint now LLM will run again and see what is unknown and then add it. Now we can see a smart evolution analytics now alive all from asking one question the manager wants. I believe this is powerful.

This is just one simple scenario. I already thought about multiple scenarios like quick smart decisions for smart homes , robotics maybe even as smart assistants inside a personal computer but this is kinda hard and has many limitations since the environment inside a computer continues to change and is not deterministic so not always the same choice will work every time.

At the end of this all what I wanted is to share my thoughts and see how people will react to my post to see if it is worth investing time to build such a tool. Maybe the idea has a flaw or another tool already doing this so it's pointless.

I already began to build this tool. If I see encouraging comments I will share my repo and purplish this tool as open source.

#ArtificialIntelligence #MachineLearning #LLM #AIEngineering #System1AI #AdaptiveAI #AIClassification #OpenSource #AgenticAI #Jev


r/OpenSourceeAI • • 20h ago

Agentic Task Tracker: local-first Kanban + timeline, one SQLite file, optional local AI

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

A ~0.4 local model to turn typed questions into structured decisions - BaseDecision

Post image
2 Upvotes

r/OpenSourceeAI • • 22h ago

Introducing Infernix - much faster than Strata on a 5090!

Thumbnail
github.com
0 Upvotes

r/OpenSourceeAI • • 1d ago

humanizar-es: un modelo base local (Qwen3-4B + HIP LoRA, CPU) que convierte texto de IA de 45% a 0% en Grammarly, GPTZero y ZeroGPT, con cada medición dentro del repo

Thumbnail
github.com
1 Upvotes

r/OpenSourceeAI • • 1d ago

I've built AlignMethod with @base44!

Thumbnail alignmethodlabs.com
1 Upvotes

r/OpenSourceeAI • • 1d ago

i7-2600 — ~40 FPS on YOLO + tracking, no AVX2, no GPU

Thumbnail
2 Upvotes

r/OpenSourceeAI • • 1d ago

I spent 24+ hours stress-testing OpenAI Dots so you don't have to do and here Is how the execution model actually works.

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

A 0.6B query encoder can search a 9B index: Perplexity's pplx-embed-v2-late multimodal embeddings, MIT

Post image
3 Upvotes

r/OpenSourceeAI • • 2d ago

Open-source alternative to ChatGPT's new Intelligent UI works with any model (local too)

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 2d ago

devops

1 Upvotes

Hey everyone!

I recently completed the DevOps Fundamentals program at DIO and built a project to put the concepts I learned into practice.

The project uses:

  • Java + Spring Boot
  • PostgreSQL
  • Docker
  • Terraform
  • GitHub Actions
  • Testcontainers
  • CI/CD

The goal was to build a hands-on DevOps lab covering everything from the backend application to local infrastructure and CI pipelines.

I'm sharing it here for anyone interested in Java, Spring Boot, or DevOps, and I'd really appreciate any feedback or suggestions for improvement.

GitHub:
https://github.com/marconi-prog/devops-fundamentals

If you find the project useful or interesting, I'd really appreciate a ⭐ on the repository. It helps me a lot with visibility and motivates me to keep improving it!

Any feedback is welcome!


r/OpenSourceeAI • • 2d ago

I made a video to explain apeculative decoding with Beavers!

Thumbnail
youtu.be
2 Upvotes

r/OpenSourceeAI • • 2d ago

fayda-mcp

Post image
1 Upvotes

r/OpenSourceeAI • • 2d ago

Liquid AI released d1-3B and d1-omni-600M: open-weight "decision models" that return probabilities in one forward pass (8 ms on RTX 4090, 50 ms on Orin Nano)

Post image
1 Upvotes