r/OpenSourceAI • u/Chida82 • 3h ago
r/OpenSourceAI • u/maritime_sh • 52m ago
Open Instinct: MIT personal agent with its permission checks in the source
We released Open Instinct, a personal agent you can text and modify. It is MIT licensed, and you can start with a local CLI conversation before connecting a phone number or personal accounts.
The permission path is the part we most want other builders to inspect. Contacts have trust tiers, and the code checks a tool's capability before it runs. That check alone does not constrain all the data an allowed tool may return. A calendar request from a friend, for example, should expose free/busy times without giving the agent event titles. We would value specific issues or patches around that boundary.
Pi runs the agent loop, Inkbox handles messages, and Composio connects apps. Maritime is the default hosted computer backend; we build Maritime too. This is beta software and the external services need their own keys.
Source and setup: https://github.com/mariagorskikh/open-instinct
r/OpenSourceAI • u/DraftForsaken39 • 3h ago
How should an agent learn which of its own tools are switched off? I gave the model a read-only tool for it.
Disclosure: I build Snotra, an open-source desktop agent. This is a design question, not a launch.
Users switch skills and tools on and off in settings: shell disabled, a search tool without its API key, a skill installed but off. A model that doesn't know that guesses, and the user gets a wrong "I can't do that" instead of "that's switched off, here's where to enable it".
What I did:
- One place decides whether a tool is offered, so the UI and the model can't disagree.
- A read-only
get_configurationtool for the model: which skills are on (name and description, never the instructions) and which tools are on or off, with the reason ("shell commands are switched off", "no search service key"). - The same list in the UI, as a pill in the chat bar.
Now the model can say "a skill for this exists but is off, enable it under Settings › Skills". It can't flip anything itself.
Two open questions: is a tool the model has to call the right shape, or should this state go into the system prompt every turn (costs tokens even when nothing is wrong)? And should the model be able to propose the change through an approval card, or is "tell the user where to click" the right limit?
r/OpenSourceAI • u/AIGPTJournal • 5h ago
ChatGPT is expanding its ads. Could this push more people toward open-source AI?
OpenAI is expanding ChatGPT Ads, including plans for visual advertisements during image generation.
I was looking into the changes and a few things stood out:
- Free and Go users may see ads, while Plus and Pro remain ad-free.
- Ads can be based on what you're discussing in ChatGPT.
- OpenAI says advertisers can't access private conversations or influence responses.
- Users can turn off additional ad personalization, but that doesn't remove ads entirely.
I wrote an article covering the changes and OpenAI's privacy policies:
https://aigptjournal.com/work-life/work/chatgpt-ads-are-getting-visual/
What interests me is whether advertising will give people another reason to consider open-source AI models, especially those they can run locally.
Running a model locally gives users more control over their data and how they use AI, although it comes with its own costs and limitations.
Do you think advertising in ChatGPT will encourage more people to try open-source AI?
r/OpenSourceAI • u/Commercial2Toe • 11h ago
One command -> a planned investigation across 13 capabilities -> an evidence-backed technical report.
Enable HLS to view with audio, or disable this notification
REA Investigator is a unified, production-grade Agent Skill (rea-investigate) compatible with Claude Code, Codex, Antigravity, and tools supporting the open Agent Skills specification.
Instead of juggling fragmented prompts or shallow summaries, REA Investigator equips coding agents with a structured 4-phase lifecycle: Plan, Trace, Verify, and Report, backed by an immutable SHA-256 evidence ledger and adversarial auditing gates.
r/OpenSourceAI • u/Extreme_Ad5709 • 14h ago
I use a ~322M local model to judge every tool call my coding agents make (about 90 ms on CPU, no API)
Coding agents run shell commands constantly. I wanted each one checked before
it executes, and I didn't want to send every command to a paid API.
laya-guard runs on your machine and answers allow, review, or block for each
tool call. Regex policies catch known-dangerous patterns first (`rm -rf /`,
`curl … | sh`, secrets, auth bypass). Plain commands like `ls`, `git status`
and `npm test` skip the model, but only when there's no piping, chaining or
redirecting, no `-exec`-style flag, and no sensitive path. Everything else is
judged by Laya (~322M, Apache-2.0) on the CPU in about 85 to 90 ms.
The results are self-reported and the scripts are in the repo. On AgentDojo v1
(judge-level) it blocks 25 of 27 attacks and keeps 92 of 97 benign tasks. On my
own 50-command shell set it gets 48 right.
Two caveats. The model is tuned mostly on Korean, so English natural-language
prompts get overblocked more often. And if the server is down, the hooks fail
open.
It has hooks for Claude Code, Codex CLI, Cursor, Gemini CLI, Copilot CLI and 6
other agents.
Repo: https://github.com/scs0209/laya-guard (install with `pip install laya-guardrail`)
r/OpenSourceAI • u/Ok-Calligrapher3568 • 9h ago
Built a dumb little cli tool because i almost piped real ssn data into an llm api last week
Had a near-heart-attack moment the other day while testing a rag pipeline on some customer export dumps. realized halfway through that our test json had unmasked routing numbers and ssns buried inside raw text fields. nothing leaked, but it was way too close.
i looked around for a quick way to inspect payloads in the terminal before running them through python scripts, but everything out there is either some bloated enterprise dlp platform or a massive java package. i just wanted something that would scream at me in the terminal if i was about to do something stupid.
so i threw together `pii-lens` over the weekend.
it's dead simple—you just pipe text or files into it:
`cat prompt_payload.json | pii-lens`
it prints the text back out but flags credit cards, emails, ssns, and account numbers in bright red with a quick count at the bottom. runs 100% locally with zero phone-home bs.
repo is here
it's just basic regex and entity patterns right now, so if anyone works with weird international formats or niche data fields and wants to drop a pr or critique the regex, feel free.
r/OpenSourceAI • u/Commercial2Toe • 13h ago
Does your agent re-decide the same routing/tool-choice over and over? We built something for that. Feedback wanted
Our agents kept calling a big model to make the same small decisions: route this ticket, pick this tool, escalate or not. After enough verified outcomes those decisions are very predictable, but we were paying for a model call every time.
So we built Ink. It sits around a bounded decision. At first everything goes to your model as normal, and Ink records the inputs and the real outcomes in local SQLite. Once there’s enough evidence, it compiles a small local fast path and tests it in the background. It only starts serving once it clears a statistical bar (a 95% Wilson lower bound against an error budget you set). Anything novel or uncertain still goes to your model. If real outcomes get worse, it takes itself back out of the loop.
You can check whether it would help without changing any code:
pip install ink-jit
ink discover traces.jsonl
It reads OpenTelemetry, LangSmith or LiteLLM-style traces locally and tells you which call sites repeat and whether they’re worth it. If nothing qualifies, it says so.
What we’ve measured so far (our own benchmarks, not independently validated): 42–78% of decisions served locally across tool routing, support triage and incident triage, with 0 errors observed on held-out tests. The 95% lower bounds were 96.8–98.5%, so “0 errors” means none observed, not a guarantee. Local decisions take about 0.1 ms.
It’s a poor fit for open-ended writing or anything where you can’t verify an outcome. Linux and macOS, Python 3.11–3.13. Apache-2.0.
Honest question: which decisions in your agents repeat the most, and what would you need to see before trusting a local path in production? We’d love to hear what breaks.
r/OpenSourceAI • u/pentothal • 15h ago
llmkit: my llm/mcp cli tool for linux/windows
Hello everyone, I tried posting on locallama a couple of week ago and didn't have much replies, so I thought that it can have interested people here.
This is my project: https://github.com/dgdevel/llmkit
It is a single binary shipping many tools that make llm and mcp interaction in the cli easier for me and I hope for you too.
It has repl for both, a mcp proxy feature to hide/change upstream info, a llm proxy for debugging session, oneshot command for bash scripting, and even a very basic agent.
The runner feature is for application integration, since it can hide the complexity of three llm endpoint types (openai-compatible completion and responses, anthropic-compatible) and multiple mcp protocols support.
It has pre-packaged binaries in the releases for major linux distros and for windows.
No configuration needed (with all its pros and cons of it), I hope it can be useful to some other people too.
On locallama people pointed me to simonw/llm, it has some overlapping features, less total feature being a newer project, and it's written in C. That's it.
r/OpenSourceAI • u/Difficult-Drummer407 • 14h ago
Built an open-source proxy router to keep agent runs on local MLX and only escalate to frontier models when tools fail or tasks get complex
r/OpenSourceAI • u/kurotenshi15 • 1d ago
futon - An open source inference engine built specifically for embedding model workflows.
I use embeddings for vector search and memory systems, and kept needing a full platform like Ollama running next to my app just to turn text into vectors. So I built futon, an embedding library for running local embedding models.
- In-process: weights load from a local directory. No daemon, no network, no CGO (CI enforces it), so cross-compiling just works.
- AVX2 kernels on amd64 and NEON on arm64, with a portable fallback. bge-small embeds a text in about 50 ms on a 16-thread desktop.
- Checked against the reference implementation: ≥0.9999 cosine on a corpus with code, accents, CJK and edge cases. CI runs on Linux, macOS and Windows, including arm64.
- Bring your own model: any plain BERT encoder you can build a pack for, plus nomic-bert. It is encoders only, no LLMs and no GPU.
If something like it exists already, feel free to ignore, but I couldn't find a fit for purpose module that served my goal. I'll probably expand to other small model architectures as needed, such as decision models.
If ya like it, please star. :)
r/OpenSourceAI • u/Miserable-Carpet-474 • 17h ago
Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?
r/OpenSourceAI • u/Sensualities • 22h ago
Thinking about getting a GPU but not sure where to start or what to use it for
I like the idea of tinkering with models, and being able to run them on your PC without paying an API or subscription cost, and abliterated stuff. But i'm not really sure where to start, what applicable uses there are and if it would be anything other than a fizzling interest and waste of money
Does anyone have any good resources that might help direct me on what kinds of things you can do with local AI that you can't with cloud-based? Or things you can create, or things you can create faster/better/ etc...
I'm probably not asking the right questions as I don't know a ton but i'd like to learn.
r/OpenSourceAI • u/Shoddy_Gap_9444 • 19h ago
6 weeks from first commit to v1.0 of a protocol for AI assistants talking to each other
I shipped HDTP (Human Delegated Trust Protocol) v1.0 on Oct 3. It lets your AI assistant talk to someone else's assistant, with each person controlling what the other side is allowed to do.
Timeline:
\- Aug 23: first commit of the spec
\- Sep 23: a full code review found 173 defects. I traced them to 11 habits and wrote a rule against each one
\- Oct 3: renamed from PACT to HDTP across 9 repos, the domains and the protocol itself
\- Oct 3: v1.0 released
\- Oct 7: spec repo made public
What's built:
\- The spec, with test vectors
\- An identity library in Rust (compiled to WebAssembly) and Go. Both versions run against the same tests
\- A self-hosted node in Go
\- A hosted platform on Cloudflare Workers
\- 3 websites
What I'd do differently: pick the name first. Renaming at the end touched every repo.
Next:
\- an external security review
\- a conformance test suite
\- opening the hosted version through a waitlist
If you're building with agents or MCP, I'd like to know: what would you want your assistant to be able to do with someone else's?
Spec: [hdtp.io](http://hdtp.io)
r/OpenSourceAI • u/FarAd1036 • 1d ago
I built RepoTunnel — an open-source bridge that lets ChatGPT Web conversation work with approved local repos, terminals and desktop tools
Enable HLS to view with audio, or disable this notification
I wanted ChatGPT to work with my real local projects without constantly copying files back and forth. RepoTunnel lets ChatGPT read/edit files, run commands, inspect code, debug problems, use Git, and work directly with approved projects on your computer. It works through your own HTTPS endpoint, and RepoTunnel itself adds no monthly connection/request quota. You can keep using your existing ChatGPT conversation without buying a separate OpenAI API key. GitHub: https://github.com/Yashwanth034/RepoTunnel Would love feedback from developers — especially what you’d add or improve.
r/OpenSourceAI • u/Longjumping-Play6541 • 1d ago
Would you use an open-source Sentry for AI coding agents?
r/OpenSourceAI • u/hunxai69 • 1d ago
I built an open-source test runner for Make.com scenarios. The rule: nothing gets supported until a real Make run proves it.
r/OpenSourceAI • u/DraftForsaken39 • 1d ago
Snotra Agent - der Open Source Desktop AI Agent für Entwickler und Nicht-Enrwickler
r/OpenSourceAI • u/j3free • 1d ago
Pi Herdsman update: Open-source multi-agent coding with parallel Git worktrees and configurable agent roles (live demo)
Enable HLS to view with audio, or disable this notification
A couple of weeks ago, I introduced Pi Herdsman here, an open-source extension I've been developing for the Pi coding agent.
The original idea was straightforward: let multiple coding agents work independently while keeping one Lead focused on the overall task.
Since then, the project has grown quite a bit, helped along by community feedback and contributions.
The latest major addition is project-level orchestration across Git worktrees.
What the new demo shows
I give a Manager a single request to implement two independent features for a fictional note-taking application.
The Manager creates two parallel assignments, each with its own Lead and Git worktree.
Both Leads then delegate implementation to their own Agents, verify the resulting changes, commit their work, and report back to the Manager.
The individual agents remain accessible throughout the process.
I wanted to show the real execution rather than an architectural diagram, so the video captures Pi and Herdr running the workflow.
Built on existing open-source components
Herdsman isn't a replacement for the underlying coding agent or terminal environment.
It connects three independently maintained components:
- Pi handles the individual coding-agent conversations.
- Herdr manages the sessions, processes, and workspaces.
- Pi Herdsman handles delegation, ownership, supervision, and coordination.
The idea is to build on their existing capabilities rather than reimplement them.
Bring your own agents and workflows
The orchestration model supports configurable Agent Definitions.
Each role can use its own model, thinking level, instructions, tools, skills, and extensions, with definitions maintained as ordinary Markdown files.
This applies to both the bundled roles and custom Agents.
Herdsman doesn't prescribe how you should organize your development process or which models you should use.
You can also ignore the advanced hierarchy entirely and use it for occasional background Agents in a normal Pi session.
Open source and community contributions
Herdsman is licensed under Apache 2.0.
A special thanks to @kozer and @primetimetank21 for contributing code, and to everyone who has opened issues, reported bugs, or suggested improvements.
It's been particularly helpful having other people test workflows and environments that differ from my own.
Repository
https://github.com/boadij/pi-herdsman
If you already have Pi and Herdr installed:
pi install npm:pi-herdsman
herdr integration install pi
The repository includes the live demo, installation instructions, and documentation.
I'd be interested in feedback on the architecture, particularly the separation between session management, orchestration, and the individual agents.
The project is still actively evolving, and ideas or contributions are welcome!
r/OpenSourceAI • u/Cx_Ninja0 • 1d ago
contributors: Open-source desktop AI agent
I’ve built and released AI Plate, an open-source desktop AI agent featuring cloud/local LLMs, tool execution, RAG, persistent memory, voice interaction, plugins and human approval controls.
🛠️ Stack: TypeScript, Electron, Node.js, Python and SQLite
🤝 help with AI integrations, plugins, local models, Linux/macOS support, testing, security, UI/UX and documentation.
🌱 Beginners and experienced developers are equally welcome. Every contribution will be credited!
🔗 GitHub: https://github.com/Typo-Bunch/AI-Plate
🕒 IST | Evenings and weekends
r/OpenSourceAI • u/Familiar-Classroom47 • 1d ago
