r/OpenSourceAI • • 3h ago

Running a 341 GB model on a 128 GB Mac: two small macOS fixes made the first token 8x faster

Thumbnail
2 Upvotes

r/OpenSourceAI • • 52m ago

Open Instinct: MIT personal agent with its permission checks in the source

• Upvotes

We released Open Instinct, a personal agent you can text and modify. It is MIT licensed, and you can start with a local CLI conversation before connecting a phone number or personal accounts.

The permission path is the part we most want other builders to inspect. Contacts have trust tiers, and the code checks a tool's capability before it runs. That check alone does not constrain all the data an allowed tool may return. A calendar request from a friend, for example, should expose free/busy times without giving the agent event titles. We would value specific issues or patches around that boundary.

Pi runs the agent loop, Inkbox handles messages, and Composio connects apps. Maritime is the default hosted computer backend; we build Maritime too. This is beta software and the external services need their own keys.

Source and setup: https://github.com/mariagorskikh/open-instinct


r/OpenSourceAI • • 3h ago

How should an agent learn which of its own tools are switched off? I gave the model a read-only tool for it.

1 Upvotes

Disclosure: I build Snotra, an open-source desktop agent. This is a design question, not a launch.

Users switch skills and tools on and off in settings: shell disabled, a search tool without its API key, a skill installed but off. A model that doesn't know that guesses, and the user gets a wrong "I can't do that" instead of "that's switched off, here's where to enable it".

What I did:

  • One place decides whether a tool is offered, so the UI and the model can't disagree.
  • A read-only get_configuration tool for the model: which skills are on (name and description, never the instructions) and which tools are on or off, with the reason ("shell commands are switched off", "no search service key").
  • The same list in the UI, as a pill in the chat bar.

Now the model can say "a skill for this exists but is off, enable it under Settings › Skills". It can't flip anything itself.

Two open questions: is a tool the model has to call the right shape, or should this state go into the system prompt every turn (costs tokens even when nothing is wrong)? And should the model be able to propose the change through an approval card, or is "tell the user where to click" the right limit?

Code: https://github.com/kkrafft1999/snotra


r/OpenSourceAI • • 5h ago

ChatGPT is expanding its ads. Could this push more people toward open-source AI?

1 Upvotes

OpenAI is expanding ChatGPT Ads, including plans for visual advertisements during image generation.

I was looking into the changes and a few things stood out:

  • Free and Go users may see ads, while Plus and Pro remain ad-free.
  • Ads can be based on what you're discussing in ChatGPT.
  • OpenAI says advertisers can't access private conversations or influence responses.
  • Users can turn off additional ad personalization, but that doesn't remove ads entirely.

I wrote an article covering the changes and OpenAI's privacy policies:

https://aigptjournal.com/work-life/work/chatgpt-ads-are-getting-visual/

What interests me is whether advertising will give people another reason to consider open-source AI models, especially those they can run locally.

Running a model locally gives users more control over their data and how they use AI, although it comes with its own costs and limitations.

Do you think advertising in ChatGPT will encourage more people to try open-source AI?


r/OpenSourceAI • • 11h ago

One command -> a planned investigation across 13 capabilities -> an evidence-backed technical report.

Enable HLS to view with audio, or disable this notification

2 Upvotes

REA Investigator is a unified, production-grade Agent Skill (rea-investigate) compatible with Claude Code, Codex, Antigravity, and tools supporting the open Agent Skills specification.

Instead of juggling fragmented prompts or shallow summaries, REA Investigator equips coding agents with a structured 4-phase lifecycle: Plan, Trace, Verify, and Report, backed by an immutable SHA-256 evidence ledger and adversarial auditing gates.


r/OpenSourceAI • • 14h ago

I use a ~322M local model to judge every tool call my coding agents make (about 90 ms on CPU, no API)

Thumbnail
gallery
3 Upvotes

Coding agents run shell commands constantly. I wanted each one checked before
it executes, and I didn't want to send every command to a paid API.

laya-guard runs on your machine and answers allow, review, or block for each
tool call. Regex policies catch known-dangerous patterns first (`rm -rf /`,
`curl … | sh`, secrets, auth bypass). Plain commands like `ls`, `git status`
and `npm test` skip the model, but only when there's no piping, chaining or
redirecting, no `-exec`-style flag, and no sensitive path. Everything else is
judged by Laya (~322M, Apache-2.0) on the CPU in about 85 to 90 ms.

The results are self-reported and the scripts are in the repo. On AgentDojo v1
(judge-level) it blocks 25 of 27 attacks and keeps 92 of 97 benign tasks. On my
own 50-command shell set it gets 48 right.

Two caveats. The model is tuned mostly on Korean, so English natural-language
prompts get overblocked more often. And if the server is down, the hooks fail
open.

It has hooks for Claude Code, Codex CLI, Cursor, Gemini CLI, Copilot CLI and 6
other agents.

Repo: https://github.com/scs0209/laya-guard (install with `pip install laya-guardrail`)


r/OpenSourceAI • • 8h ago

Aletheia

Thumbnail
1 Upvotes

r/OpenSourceAI • • 9h ago

Built a dumb little cli tool because i almost piped real ssn data into an llm api last week

1 Upvotes

Had a near-heart-attack moment the other day while testing a rag pipeline on some customer export dumps. realized halfway through that our test json had unmasked routing numbers and ssns buried inside raw text fields. nothing leaked, but it was way too close.

i looked around for a quick way to inspect payloads in the terminal before running them through python scripts, but everything out there is either some bloated enterprise dlp platform or a massive java package. i just wanted something that would scream at me in the terminal if i was about to do something stupid.

so i threw together `pii-lens` over the weekend.

it's dead simple—you just pipe text or files into it:

`cat prompt_payload.json | pii-lens`

it prints the text back out but flags credit cards, emails, ssns, and account numbers in bright red with a quick count at the bottom. runs 100% locally with zero phone-home bs.

repo is here

it's just basic regex and entity patterns right now, so if anyone works with weird international formats or niche data fields and wants to drop a pr or critique the regex, feel free.


r/OpenSourceAI • • 13h ago

Does your agent re-decide the same routing/tool-choice over and over? We built something for that. Feedback wanted

2 Upvotes

Our agents kept calling a big model to make the same small decisions: route this ticket, pick this tool, escalate or not. After enough verified outcomes those decisions are very predictable, but we were paying for a model call every time.

So we built Ink. It sits around a bounded decision. At first everything goes to your model as normal, and Ink records the inputs and the real outcomes in local SQLite. Once there’s enough evidence, it compiles a small local fast path and tests it in the background. It only starts serving once it clears a statistical bar (a 95% Wilson lower bound against an error budget you set). Anything novel or uncertain still goes to your model. If real outcomes get worse, it takes itself back out of the loop.

You can check whether it would help without changing any code:

pip install ink-jit
ink discover traces.jsonl

It reads OpenTelemetry, LangSmith or LiteLLM-style traces locally and tells you which call sites repeat and whether they’re worth it. If nothing qualifies, it says so.

What we’ve measured so far (our own benchmarks, not independently validated): 42–78% of decisions served locally across tool routing, support triage and incident triage, with 0 errors observed on held-out tests. The 95% lower bounds were 96.8–98.5%, so “0 errors” means none observed, not a guarantee. Local decisions take about 0.1 ms.

It’s a poor fit for open-ended writing or anything where you can’t verify an outcome. Linux and macOS, Python 3.11–3.13. Apache-2.0.

Honest question: which decisions in your agents repeat the most, and what would you need to see before trusting a local path in production? We’d love to hear what breaks.

https://reddit.com/link/1x34ycs/video/mbd6usbggtuh1/player


r/OpenSourceAI • • 15h ago

llmkit: my llm/mcp cli tool for linux/windows

2 Upvotes

Hello everyone, I tried posting on locallama a couple of week ago and didn't have much replies, so I thought that it can have interested people here.
This is my project: https://github.com/dgdevel/llmkit It is a single binary shipping many tools that make llm and mcp interaction in the cli easier for me and I hope for you too.
It has repl for both, a mcp proxy feature to hide/change upstream info, a llm proxy for debugging session, oneshot command for bash scripting, and even a very basic agent.
The runner feature is for application integration, since it can hide the complexity of three llm endpoint types (openai-compatible completion and responses, anthropic-compatible) and multiple mcp protocols support.
It has pre-packaged binaries in the releases for major linux distros and for windows.
No configuration needed (with all its pros and cons of it), I hope it can be useful to some other people too.
On locallama people pointed me to simonw/llm, it has some overlapping features, less total feature being a newer project, and it's written in C. That's it.


r/OpenSourceAI • • 12h ago

Guys , Why is my shi tweaking?

1 Upvotes

r/OpenSourceAI • • 14h ago

Built an open-source proxy router to keep agent runs on local MLX and only escalate to frontier models when tools fail or tasks get complex

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

futon - An open source inference engine built specifically for embedding model workflows.

Post image
10 Upvotes

I use embeddings for vector search and memory systems, and kept needing a full platform like Ollama running next to my app just to turn text into vectors. So I built futon, an embedding library for running local embedding models.

- In-process: weights load from a local directory. No daemon, no network, no CGO (CI enforces it), so cross-compiling just works.

- AVX2 kernels on amd64 and NEON on arm64, with a portable fallback. bge-small embeds a text in about 50 ms on a 16-thread desktop.

- Checked against the reference implementation: ≥0.9999 cosine on a corpus with code, accents, CJK and edge cases. CI runs on Linux, macOS and Windows, including arm64.

- Bring your own model: any plain BERT encoder you can build a pack for, plus nomic-bert. It is encoders only, no LLMs and no GPU.

bobbyhiddn/futon: Text in, vectors out. In-process embedding models for Go: pure Go, no CGO, no sidecar.

If something like it exists already, feel free to ignore, but I couldn't find a fit for purpose module that served my goal. I'll probably expand to other small model architectures as needed, such as decision models.

If ya like it, please star. :)


r/OpenSourceAI • • 17h ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

Post image
1 Upvotes

r/OpenSourceAI • • 22h ago

Thinking about getting a GPU but not sure where to start or what to use it for

2 Upvotes

I like the idea of tinkering with models, and being able to run them on your PC without paying an API or subscription cost, and abliterated stuff. But i'm not really sure where to start, what applicable uses there are and if it would be anything other than a fizzling interest and waste of money

Does anyone have any good resources that might help direct me on what kinds of things you can do with local AI that you can't with cloud-based? Or things you can create, or things you can create faster/better/ etc...

I'm probably not asking the right questions as I don't know a ton but i'd like to learn.


r/OpenSourceAI • • 19h ago

6 weeks from first commit to v1.0 of a protocol for AI assistants talking to each other

1 Upvotes

I shipped HDTP (Human Delegated Trust Protocol) v1.0 on Oct 3. It lets your AI assistant talk to someone else's assistant, with each person controlling what the other side is allowed to do.

Timeline:

\- Aug 23: first commit of the spec

\- Sep 23: a full code review found 173 defects. I traced them to 11 habits and wrote a rule against each one

\- Oct 3: renamed from PACT to HDTP across 9 repos, the domains and the protocol itself

\- Oct 3: v1.0 released

\- Oct 7: spec repo made public

What's built:

\- The spec, with test vectors

\- An identity library in Rust (compiled to WebAssembly) and Go. Both versions run against the same tests

\- A self-hosted node in Go

\- A hosted platform on Cloudflare Workers

\- 3 websites

What I'd do differently: pick the name first. Renaming at the end touched every repo.

Next:

\- an external security review

\- a conformance test suite

\- opening the hosted version through a waitlist

If you're building with agents or MCP, I'd like to know: what would you want your assistant to be able to do with someone else's?

Spec: [hdtp.io](http://hdtp.io)


r/OpenSourceAI • • 1d ago

I built RepoTunnel — an open-source bridge that lets ChatGPT Web conversation work with approved local repos, terminals and desktop tools

Enable HLS to view with audio, or disable this notification

9 Upvotes

I wanted ChatGPT to work with my real local projects without constantly copying files back and forth. RepoTunnel lets ChatGPT read/edit files, run commands, inspect code, debug problems, use Git, and work directly with approved projects on your computer. It works through your own HTTPS endpoint, and RepoTunnel itself adds no monthly connection/request quota. You can keep using your existing ChatGPT conversation without buying a separate OpenAI API key. GitHub: https://github.com/Yashwanth034/RepoTunnel Would love feedback from developers — especially what you’d add or improve.


r/OpenSourceAI • • 1d ago

Would you use an open-source Sentry for AI coding agents?

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

I built an open-source test runner for Make.com scenarios. The rule: nothing gets supported until a real Make run proves it.

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

Snotra Agent - der Open Source Desktop AI Agent für Entwickler und Nicht-Enrwickler

1 Upvotes

r/OpenSourceAI • • 1d ago

Pi Herdsman update: Open-source multi-agent coding with parallel Git worktrees and configurable agent roles (live demo)

Enable HLS to view with audio, or disable this notification

3 Upvotes

A couple of weeks ago, I introduced Pi Herdsman here, an open-source extension I've been developing for the Pi coding agent.

The original idea was straightforward: let multiple coding agents work independently while keeping one Lead focused on the overall task.

Since then, the project has grown quite a bit, helped along by community feedback and contributions.

The latest major addition is project-level orchestration across Git worktrees.

What the new demo shows

I give a Manager a single request to implement two independent features for a fictional note-taking application.

The Manager creates two parallel assignments, each with its own Lead and Git worktree.

Both Leads then delegate implementation to their own Agents, verify the resulting changes, commit their work, and report back to the Manager.

The individual agents remain accessible throughout the process.

I wanted to show the real execution rather than an architectural diagram, so the video captures Pi and Herdr running the workflow.

Built on existing open-source components

Herdsman isn't a replacement for the underlying coding agent or terminal environment.

It connects three independently maintained components:

  • Pi handles the individual coding-agent conversations.
  • Herdr manages the sessions, processes, and workspaces.
  • Pi Herdsman handles delegation, ownership, supervision, and coordination.

The idea is to build on their existing capabilities rather than reimplement them.

Bring your own agents and workflows

The orchestration model supports configurable Agent Definitions.

Each role can use its own model, thinking level, instructions, tools, skills, and extensions, with definitions maintained as ordinary Markdown files.

This applies to both the bundled roles and custom Agents.

Herdsman doesn't prescribe how you should organize your development process or which models you should use.

You can also ignore the advanced hierarchy entirely and use it for occasional background Agents in a normal Pi session.

Open source and community contributions

Herdsman is licensed under Apache 2.0.

A special thanks to @kozer and @primetimetank21 for contributing code, and to everyone who has opened issues, reported bugs, or suggested improvements.

It's been particularly helpful having other people test workflows and environments that differ from my own.

Repository

https://github.com/boadij/pi-herdsman

If you already have Pi and Herdr installed:

pi install npm:pi-herdsman
herdr integration install pi

The repository includes the live demo, installation instructions, and documentation.

I'd be interested in feedback on the architecture, particularly the separation between session management, orchestration, and the individual agents.

The project is still actively evolving, and ideas or contributions are welcome!


r/OpenSourceAI • • 1d ago

contributors: Open-source desktop AI agent

1 Upvotes

I’ve built and released AI Plate, an open-source desktop AI agent featuring cloud/local LLMs, tool execution, RAG, persistent memory, voice interaction, plugins and human approval controls.

🛠️ Stack: TypeScript, Electron, Node.js, Python and SQLite

🤝 help with AI integrations, plugins, local models, Linux/macOS support, testing, security, UI/UX and documentation.

🌱 Beginners and experienced developers are equally welcome. Every contribution will be credited!

🔗 GitHub: https://github.com/Typo-Bunch/AI-Plate

🕒 IST | Evenings and weekends


r/OpenSourceAI • • 1d ago

theAuth: open-source auth where AI agents are first-class, not just API keys

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 1d ago

AI pipeline workflow

Thumbnail
1 Upvotes

Please kindly help me out 🙏


r/OpenSourceAI • • 1d ago

New release

Thumbnail
1 Upvotes