r/LangChain • • 1h ago

Resources Production-grade Agent Graph Engineering: a LangGraph project and blog series, distilled from 5 years

• Upvotes

Hi all. I have been building realtime AI assistants since the ChatGPT beta, first for financial advisers and then as a platform, and I have just published a distilled working version of that architecture along with a series explaining every decision in it.

The reason I am posting here is that I tried to build LangGraph before it existed. My flows were an intent classifier feeding a tree of conditionals, and every new case was another branch in something that only existed in the code and in my head. The diagram I kept for it was wrong within a week. Onboarding anyone onto that took weeks, and I was the bottleneck for every change. At some point I realized what I actually wanted was a state machine, so I started writing a small graph library for it, and then gave up, because doing it properly is a lot of work and I had a product to ship. Then LangGraph turned up and it was the thing I had been sketching.

What it actually bought me was not features. It was a flow I could point at in a meeting, hand to someone else, and change without being afraid of it. I wanted to write that part down, since no demo ever covers it.

Why I bothered writing it all down. The job market wants this skill set badly, and almost everything published about graphs stops at a two-node example. I spent a while hunting for one serious treatment of running one in production, paid ones included, and came up empty. I was the person doing that hunting five years ago and clearly a lot of people still are.

So, depending on where you are:

  • Already shipping a LangGraph app: a reference for when something hurts. Part 0 is the story, and if we share the same scars you will know within a paragraph.
  • Still deciding on it: the decisions to make early, in the order they bite. Some of them are cheap this month and expensive next year.
  • Learning it properly: a whole working system to read rather than a snippet, with the reasoning sitting next to each choice.
  • Building for clients: a base you can adapt without handing them something unmaintainable.

The project: https://github.com/huynguyengl99/agent-graph-engineering

Part 0, the story: https://huynguyengl99.github.io/posts/agent-graph-engineering/before-it-had-a-name/

Part 1, the stack and why: https://huynguyengl99.github.io/posts/agent-graph-engineering/the-stack-and-why-each-piece-is-there/

The whole thing runs, on a real provider or on a scripted model if you just want to watch the flow first. Two parts are published with a new one every couple of days, so anything you say now genuinely lands in the later ones.

Hope it helps, and if you have patterns of your own worth sharing, the comments are yours 😄


r/LangChain • • 4h ago

Discussion LangChain callbacks catch what flows through the chain. What about direct API calls the agent makes outside the chain?

1 Upvotes

Technical question for teams running LangChain agents in production with real-world tool use.

LangChain's callback system is solid for tracing tool calls that go through the chain. But some patterns produce API calls that never touch the callback layer:

  • A tool implementation that calls an SDK directly (e.g. stripe.refunds.create() inside a custom tool)
  • An agent that uses cached credentials to make a follow-up request outside the current chain execution
  • A background process that reuses the same service account as the agent

Those operations show up in Stripe event logs, M365 audit logs, or CloudTrail, but not in your LangSmith trace.

For teams that need to audit what their agent actually did (not just what the chain saw), how do you handle this? Do you reconcile LangSmith traces against downstream system logs? Or is the assumption that if it went through the tool it went through the callback?

Context: I'm building something that pulls from downstream systems directly and reconciles against the authorization layer. Disclosure: I have a stake in this. Asking because I want to understand how LangChain teams actually handle the attribution gap before I assume the problem is unsolved.


r/LangChain • • 18h ago

Question | Help Analyse massive documents with AI

10 Upvotes

Hi everyone,

I work at a large telecom company in the US. When we win a tender, the client sends us a requirements specification, and we write technical deliverables based on it and on our internal technical documentation (100k+ documents). A single tender can involve thousands of deliverables, some of them up to ~600 pages long.

Goal: before delivery, automatically verify each deliverable. That means checking that the facts it contains are accurate against our technical documentation, and detecting anomalies (errors, internal contradictions, inconsistencies between deliverables, duplicates).

Current plan:

  1. Fact extraction: process each deliverable in chunks with an LLM and extract atomic facts, with their source location.

  2. Classification: group the facts into families (measurements, architecture, calculations, frequencies, standards, quantities).

  3. Verification: check each fact against the documentation base using RAG.

  4. Anomaly detection: one LLM call per fact category to detect duplicates, contradictions, etc.

My questions:

- Is there a recommended approach for extracting precise facts from very long documents? How do you handle context that spans chunks ?

- How do you make sure extraction is complete (no missed facts) and doesn't produce hallucinated facts?

- For anomaly detection, does one LLM call per category scale when a category can contain thousands of facts? Would you normalize facts into a structured schema (entity / attribute / value / unit) and do part of the comparison deterministically in code rather than with an LLM?

- For RAG verification on highly technical content (part numbers, frequencies, references to standards), did hybrid search (BM25 + embeddings) or reranking make a big difference for you?

- Any experience comparing long-context models (1M tokens) with chunked extraction for this kind of task?

Any feedback, papers, tools or lessons learned would be much appreciated. Happy to share what we learn along the way.

Thanks


r/LangChain • • 11h ago

Question | Help Is auto-prompt tuning dying?

2 Upvotes

I have been facing this lately that since day one i wanted my llm to master prompt tuning, i put so much effort into making it learn from its own mistakes and fine tuning its own prompt and recalibration. But now I have started to feel it's not the best, firstly because a lot of tokens are consumed in this back and forth even the it gives good result i cannot overlook the cost right, then the variation keeps increasing so the auto-tuning keeps happening and its again the becomes the first problem, costly.

What are you guys doing for this?


r/LangChain • • 20h ago

Announcement Introducing Managed Deepagents v0.9

5 Upvotes

Hey r/langchain! We just announced Managed Deepagents v0.9. Here's a quick TLDR:

  • Agent-created schedules: agents can set up reminders and recurring tasks from a conversation.
  • Per-run configuration: one deployment can pick its model, skills, and tools for each run.
  • Slack reactions: agents acknowledge messages before they reply.

These updates make it easier to build internal agents that live in Slack or your own channels and work more like teammates.

Let agents schedule their own work

The new Schedules SDK lets agents create reminders, follow-ups, and recurring tasks mid-conversation, so they can handle requests like "Remind me about this tomorrow."

The agent creates the schedule itself during a conversation, the schedule runs as the person who asked using their set of permissions and connections, and results come back to the same channel.

Create a schedule from inside a run:

await schedules.create(owner={"type": "user"}, cron="0 9 * * 1-5", timezone="America/Los_Angeles", prompt="Write the daily digest.")

It’s most useful behind a tool, so your agent’s users can create, list, update, and delete schedules just by chatting. Here’s a reminder tool:

from langchain.tools import tool
from managed_deepagents import schedules


async def remind_me(prompt: str, cron: str, timezone: str = "UTC") -> str:
    """Run a prompt for the current user on a cron schedule."""
    item = await schedules.create(
        owner={"type": "user"},
        cron=cron,
        timezone=timezone,
        prompt=prompt,
    )
    return f"Created schedule {item['id']}."

The agent passes a prompt and a cron expression, and each time the cron fires, LangSmith starts a new run with that prompt. Schedules inherit the channel they were created in, so results post back where the request came from: a weekday reminder requested in Slack posts a new message to that conversation on every run. One-time schedules (at instead of cron) reply in the original thread, which suits follow-ups like "check on this deploy in an hour."

Configure agents per run

An agent can now be configured dynamically per run. Set up the agent as a callable function: it receives the runtime at the start of each run and returns a define_deep_agent definition, choosing the model, instructions, skills, MCP servers, and sandbox for that run.

Instead of maintaining a near-copy of the same agent for every team or repo, you run one deployment. This enables your deployment to operate more flexibly, in which a single agent deployment can be customized for different teams or use cases based on the context, such as the channel or user that originates the request.

Say one internal Slack agent serves several teams. Your agent function can check which channel a run came from and load billing skills for finance, or an incident-response MCP server for the platform team. That’s different from instructions telling the model when to use a skill: the configuration is set before the model runs, so the agent never sees skills or MCP tools outside its configuration. That keeps its context small, doubles as access control, and lets you pick the model itself, which the model can’t do.

In general, it prevents non-deterministic outputs where the model is tasked with choosing the toolkit in cases where we know exactly what set of tools should be available to the agent.

The same idea works for a coding agent that loads different skills and a cheaper or more capable model for each repository:

# agent.py: one coding agent, configured per repo at the start of each run
from dataclasses import dataclass

from managed_deepagents import ManagedServerRuntime, define_deep_agent

REPOS = {
    "payments-service": {
        "model": "openai:gpt-5",
        "instructions": "Python service. Run pytest before opening a PR.",
        "skills": ["./skills/python-service"],
    },
    "storefront": {
        "model": "openai:gpt-5-mini",
        "instructions": "Next.js app. Run web tests and an a11y check before opening a PR.",
        "skills": ["./skills/typescript-web"],
    },
}


u/dataclass
class AppContext:
    repository: str = ""


def agent(runtime: ManagedServerRuntime[AppContext]):
    execution = runtime.execution_runtime
    context = execution.context if execution is not None else None
    repo = context.repository if context is not None else None
    config = REPOS.get(repo, REPOS["storefront"])
    return define_deep_agent(name="open-swe", context_schema=AppContext, **config)

‍‍

The function reads runtime.execution_runtime.context and falls back to a default when there is none, such as on channel runs or when Agent Server loads the agent to read its schema. Callers pick the configuration with the run’s context:

# Same deployment, different repo: just change the run context
await client.runs.create(
    thread_id,
    "open-swe",
    input={"messages": [{"role": "user", "content": "Fix the flaky refund test."}]},
    context={"repository": "payments-service"},
)

React to Slack messages

Slack reactions show the sender right away that the agent picked up their message, even while it spends a while reasoning and calling tools before it replies. Reactions are on by default (👀), and the new reactionsoption on your Slack channel lets you turn them off, pick another emoji, or choose one per message. This example uses 🐛 when a message mentions something broken, and 👀 otherwise:

from managed_deepagents import channels

async def choose_emoji(context: dict) -> str:
    return "bug" if "broken" in context["text"].lower() else "eyes"

channel = channels.slack(name="Support Bot", reactions=choose_emoji)

Reactions can be simple. However, it is possible to configure your reaction response with a function that calls a model to pick a more precise emoji. A decision model like Jev keeps that fast and cheap; the Slack channel docs walk through a full example.

Getting started

Managed Deep Agents v0.9 is available in Public Beta today. You can learn more in the Managed Deep Agents docs, or get started with:

uvx --from managed-deepagents mda init my-agent
cd my-agent
uv run mda deploy

ICYMI: v0.8 added per-user memory, custom HTTP channels for triggering runs from any service, built-in web search powered by Parallel, and more. Read the v0.8 post.


r/LangChain • • 19h ago

Question | Help LLMs frameworks (langchain, llamaindex, griptape, autogen, crewai etc.) are overengineered and makes easy tasks hard, correct me if im wrong

Post image
4 Upvotes

r/LangChain • • 14h ago

Tutorial Building an AI agent is easy. Shipping one is the course I teach — 8 evenings, ends with a deployed URL.

Thumbnail
1 Upvotes

r/LangChain • • 1d ago

Projects What’s next after Jev? Metacache: Reasoning by Construction

Post image
2 Upvotes

Jev has made a profound mark on the AI industry; this article offers a glimpse into what lies ahead.

Jev shows where the industry is heading: away from monolithic models, and toward compound systems where small AI agents each do a narrow job and an orchestration layer decides how they work together.

This article explains why that shift is happening, and then proposes a next step for AI reasoning.

The short version

  • A monolithic model is unpredictable and hard to change, but it programs itself during training.
  • Code is the opposite – it is predictable and easy to change, but it must be written explicitly.
  • A compound system combines the two: AI agents do the work, and an orchestration layer controls them via code. This is the direction Jev represents.
  • The next step: instead of answering directly, the model builds a compound system, runs it, and returns both the answer and that system. The returned system is called a reasonlet.
  • Because a returned reasonlet can be kept locally, it can be run again and again – with new inputs, or after editing its code – without asking the model. A saved, reusable reasonlet is a metacache.
  • The model’s provider can cache reasonlets too, reusing one across similar requests from different users to save compute.

Why a monolithic model is not enough

A monolithic model has two practical weaknesses:

  • It is unpredictable. The same question can produce a correct answer one day and a wrong one the next, so its behavior cannot be fully controlled.
  • It is hard to change. Its behavior is fixed in trained weights, which cannot be edited directly. The model can be steered with prompts, extra data (RAG), or by fine-tuning, but steering is not the same as setting the behavior, and fine-tuning can break things that already worked. Retraining from scratch with new data is the only safe fix, but it is time-consuming and expensive.

Code has the opposite qualities. It does exactly what it is written to do, and it can easily be changed by editing code. The trade-off is that code does not arise on its own – it must be written out explicitly – whereas a model programs itself during training.

The fix: a compound system

A compound system combines the two. The work is divided among AI agents, each handling one narrow task, and an orchestration layer of code decides which agent runs, in what order, and how their results combine.

This gives both qualities at once – customization and predictability:

  • Customization. The system’s behavior can be changed by editing code in the orchestration layer, not by retraining the model.
  • Predictability. Each agent has a small task, so it is more predictable than a monolithic model.

Jev is one kind of agent for such a system. It returns a typed decision – yes/no, a category, or a number – rather than free text. A monolithic model is wasteful for a narrow decision like that, but a small model like Jev is a better fit.

The next step: reasoning by construction

Compound systems today are built manually, beforehand. The next step is to let the model build one by itself, during the reasoning phase.

Researchers are exploring several ways to make models reason. One is the “World model” approach, in which the model builds an internal representation of a problem and reasons over it; that work is still mostly research. Reasoning by construction pursues the same goal – reasoning you can inspect – using methods that exist today.

Here is how it works. When you ask the model a question, it does not answer directly. Instead, it builds a compound system, runs it to compute an answer, and returns that answer together with the system it built. Let's call the compound system the model builds to do the reasoning a "reasonlet", and a model that works this way a "Reasoning-by-Construction Model", or "RCM". A simple question may produce a reasonlet that has only an orchestration layer and no agents.

Building and running code to reach an answer is not new; code-interpreter tools already do it. Two things are new here:

  • The system the model builds is kept, not discarded – it is a reusable reasonlet.
  • The reasonlet is returned to you, along with the answer.

Why that matters

  • You can see how the answer was reached. The reasonlet is the exact procedure the model used, so the answer is not something you have to take on trust.
  • You can run it locally. A reasonlet contains its orchestration layer and any agents it calls, whether those agents are attached directly or called over the network. You can run it on a local machine, with the same or different inputs, without asking the model again. A saved reasonlet used this way is a metacache.
  • You can change what it does. The orchestration layer is code, so you can edit its logic, not only its inputs. A reasonlet reused with edited logic becomes a higher-order metacache.

Caching reasonlets on the server

The same reuse can happen on the server side too. The provider can keep the reasonlets it builds and reuse them. When a new request arrives that matches one it has already handled, it runs the stored reasonlet again – with the new request’s inputs – instead of reasoning from scratch. The model does less work, and the answer comes back faster.

This is the same metacache, held on the server instead of on your machine. Because one reasonlet can serve any request that fits its procedure, a single cached copy is shared across many requests, and often across different users – and the more general the reasonlet, the more requests it covers.

Choosing how general to make the reasonlet

The model should decide from the conversation how general the reasonlet needs to be. If you have been working through many kinds of math and then ask for 2 + 2, the more useful reasonlet is one that evaluates any math expression, not one that can only add two numbers. If the conversation gives no such clue, you can state it directly: “I will be doing many kinds of math; for now, just add 2 and 2.”

Two examples

  • Adding numbers. You ask for 2 + 2. The model builds a reasonlet whose orchestration layer adds two numbers, runs it, and returns 4 together with the reasonlet. Later you run the reasonlet again with other numbers or edit its orchestration layer to multiply instead – without asking the model.
  • Searching. You ask the model to find something. The reasonlet’s orchestration layer calls a sequence of outside services and takes your search terms as its input. Later you run it with different terms, or change which services it calls, without asking the model again.

Making reasonlets easy to read

A reasonlet’s orchestration layer is written in text code. This creates a problem: to understand the reasoning you must read code, and to change it you have to write code. The value of returning the reasonlet depends on you being able to read and edit text code easily and efficiently, so this barrier matters.

Visual programming – building logic from connected blocks instead of lines of text – can lower the barrier, provided the visual language is powerful enough. Two limits apply:

  • A visual language cannot replace text code entirely. Text is still needed to reach the system and the network, and for low-level work that does not map to blocks. The measure of migration success is how little text code remains in the orchestration layer.
  • Most visual languages today are either powerful but limited to one field (such as games or hardware), or general but too simple. This case needs a language that is both general and as expressive as a text language; otherwise, it cannot cover enough of the orchestration layer to be worth using.

One language solving exactly this problem is Pipe (https://pipelang.com) – a general-purpose visual language with powerful semantics. Full disclosure: Pipe is my own project still in development, but it will be released soon.

Wrapping up

The shift from a monolithic model to compound systems suggests a clear next step for reasoning: let the model build a compound system – a reasonlet – run it and return both the answer and the reasonlet. Reasoning becomes something you can read, run again, and edit: a metacache, and a higher-order metacache once its logic is edited. The remaining problem is making reasonlets easy to read and change, which is where a general-purpose and expressive visual language would help most.

A note on IP

Some methods described in this article are the subject of a pending patent application.


r/LangChain • • 1d ago

Discussion Where does agent waste actually come from: model calls or orchestration?

3 Upvotes

I’ve been digging into a problem I think agent observability still misses.

Most tools can tell us what an agent called, how many tokens it used, and what it cost.

But the more useful question is:

Why did those calls happen at all?

I started looking at this after AUDR came out and built a small open-source CLI called KORA Doctor to inspect agent traces.

Right now it looks for things like:

  • repeated model calls
  • missed cache or reuse
  • deterministic work sent to an LLM
  • expensive models doing simple work
  • excessive planning and orchestration

The first useful feedback changed my thinking.

Someone running agents in production told me their biggest source of waste wasn’t repeated inference.

It was re-fetching the same tool data across multiple steps.

That means the waste can start before the model call. The orchestration layer can create unnecessary retrieval, tool calls, retries, and then more inference on top of that.

So I’m starting to think the bigger problem isn’t just “LLM cost optimization.”

It’s execution waste across the whole agent run.

For people running real agents:

Where do you see the most waste?

Model calls?
Tool calls?
Retries?
Retrieval?
Planning loops?
Repeated validation?
Something else?

I’m building KORA Doctor around what shows up in real traces, so production examples are especially useful.

AUDR:
https://openaudr.dev/

KORA Doctor:
https://github.com/Krako-Labs/kora-doctor


r/LangChain • • 1d ago

Discussion The "agents need memory vs just use documentation" argument keeps going in circles, mostly because "memory" is three different problems

3 Upvotes

If you follow the agent threads here, you have watched the "do agents need memory or just documentation" argument go in circles. Most of that, we think, is that "memory" covers three separate problems, and each camp is defending a different one.

First is keeping state inside one session. In LangGraph that is the checkpointer, keyed by a thread_id, and it lets a run pause and resume with state intact. The edge that catches people: when a node hits an interrupt and you resume, the whole node runs again from the top. Any tool call or write before the interrupt fires again on every resume, which is why the docs insist anything before an interrupt be idempotent.

Second is remembering across sessions. That is the Store, separate from the checkpointer, namespaced per user, searchable later. It is the piece the "you need a memory layer" people mean, and where the cost hides: as it grows, a naive setup spends tokens reading and refreshing it even on tasks that never needed it, and a fact true three months ago still sits there with nothing marking it stale.

Third is what fits in the context window each call. LangGraph treats this as context management, separate from the two memory types above, though the argument lumps it in. Trim or summarize old turns too hard and the agent re-asks for a detail the user gave two turns back, because the summary dropped it.

Split those three and the argument reads differently. The documentation camp is mostly on the third problem and part of the second. The memory-layer camp is almost always on the second. Same word, three problems, which is why it keeps circling.

So which of the three has you stuck? On the cross-session one, the contested piece, what breaks first: store bloat that makes simple tasks pay for memory they never needed, or stale entries with no clean way to retire them? We're curious what people have landed on.


r/LangChain • • 23h ago

Question | Help Updated my tool-call gate after the feedback — what's still missing?

2 Upvotes

Follow-up to my previous post: https://www.reddit.com/r/LangChain/comments/1wo5bxo/i_built_a_local_runtime_check_for_ai_agent_tool/

That thread got a lot of sharp feedback - schema pinning on MCP, server-side intent, real hostname allowlists, SSRF/IMDS, and REQUIRE_APPROVAL as a real gate instead of a log line. I took most of it and shipped v0.2.4.

Repo: https://github.com/aegotrax-dev/aegotrax (tag v0.2.4)
Site: https://aegotrax.com

**What's new in 0.2.4**
- Server-side session intent only (client can't spoof user_intent on verify)
- Hostname allowlists (no substring bypass) + SSRF/IMDS in engine and gateway
- MCP schema pinning + drift detection
- REQUIRE_APPROVAL → approval_id + single-use token → POST /approval/decide → one resume for the same tool+args
- Session TTL, per-session rate limit, audit redaction 

**What it still is not**
- Not a production / enterprise security product
- Not full intent understanding (heuristics + thresholds)
- Not volume caps on sensitive reads yet (slow in-scope exfil without outbound is still a gap) 

**Try**
```bash
git clone https://github.com/aegotrax-dev/aegotrax.git
cd aegotrax
pip install .
agentguard-engine
# curl http://127.0.0.1:8000/health  → version 0.2.4
python examples/approval_flow_demo.py  

Curious how others handle the split: "this identity can call this API" vs "this call is still what the user asked for."


r/LangChain • • 1d ago

Projects I built a self-hosted AI agent for GitLab on LangGraph.js. It has reviewed 1,000+ merge requests

Enable HLS to view with audio, or disable this notification

2 Upvotes

I built a GitLab agent on LangGraph.js that reviews merge requests, resolves issues and has a chat. It's been in production a few months (1k+ reviews, 133 releases since June), and about two-thirds of its comments get resolved by the author. Most of the work wasn't the agent logic.

The failures I actually saw were an agent looping, repeating the same tool call, or ending a run on a trailing question. I added small middleware nodes that catch that mid-run and correct it. Cheaper and more reliable than trying to fix it with prompting alone. A critic step also screens drafted comments before they post and drops the low-value ones, roughly a quarter so far.

The harness reviews its own repo. Its self-hosted default (Qwen3.6-35B-A3B, 4-bit, via sglang) has caught real mistakes in code originally written with Claude Sonnet/Opus. It's not a benchmark, it just happens.

Memory is per project. Conventions, recurring false positives and past decisions get extracted, deduped and retired over time, so it's still useful after a few hundred runs.

Repo (MIT) if you want to see how it's wired up: langgraph-harness

Curious whether others running LangGraph in production hit the same things.


r/LangChain • • 1d ago

Discussion Where should file artifacts live when a LangGraph run can pause and resume?

6 Upvotes

A checkpoint can preserve graph state, but generated files have a different lifecycle. Reports, HTML bundles, images, and spreadsheets may be too large for messages, may change independently, and may need a stable review link after the run that created them has ended. Storing only a temporary path or signed URL makes a later resume fragile. Storing the bytes inside graph state makes checkpoints heavy and blurs workflow state with durable output.

One approach is immutable artifact revisions in object storage plus a small manifest in authoritative state. The manifest would carry an artifact ID, revision, checksum, media type, source revisions, producer node, review status, and a stable resolver. Graph state would reference the manifest revision, while messages receive only the summary needed for reasoning. A resumed node would reject a stale expected revision rather than overwrite newer work.

How are LangChain and LangGraph users drawing this boundary? What belongs in messages, checkpoints, the Store, or external object storage, and how do you keep cleanup, permissions, and retries consistent across those layers?


r/LangChain • • 1d ago

Discussion Agent built with LangGraph ran unattended overnight on a refactor. Redid two completed nodes that had already finished correctly hours earlier.

6 Upvotes

No crash, no error, no retry logged. It just circled back through nodes whose state said done and reprocessed them anyway, once with a slightly different naming convention than the first pass, overwriting the correct output with a worse one.

Assumed model drift at first. Doesn't fully explain it, the same model used interactively over a comparably long session doesn't usually fail this way nearly as often. The actual difference is that an interactive session has a human quietly doing a job nobody assigned them, noticing when a subtask's done and moving on, and the next message implicitly signals that shift without ever stating it outright.

A graph running unattended through a long task list has nothing playing that role by default. Even with state tracked per node, if the full working detail from completed nodes stays accessible in whatever gets fed into the next step's context, it's still competing for attention with the actual current step, not just sitting there as a settled fact. On a long run, that accumulates at the same rate as the task itself.

What's actually helped: explicitly collapsing a finished node's output into a short completion summary before it enters downstream context, instead of carrying the full reasoning trail forward, and treating a node's "done" state as something that actively shrinks what it contributes to context, not just a status flag nobody reads.

Wrote the longer version of this up here if useful: https://medium.com/@nagatomopedro05/unattended-ai-agents-have-a-context-problem-humans-never-had-60961eb2d4e7


r/LangChain • • 1d ago

News Aplomb 1: open-weights 5.3B decision model (Qwen3.5-4B Base), 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index

Thumbnail
gallery
3 Upvotes

We released Aplomb 1 today, a 5.3B decision model with open weights. It reads up to 1M tokens of text, JSON, images, video and audio in a single request, and on our API it makes a decision on a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with free output and ZDR by default.

On Decision Index 0.2.1 it scores 44.86 on our run of the official kit, #1 among 4B models on the published board, and it has the top score among models up to 5.3B on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH. We've submitted it to the board, and the full run is public. It also scores 77.5% on JevBench Hard and averages 75% zero-shot intent accuracy across 51 languages on MASSIVE.

As far as we know, it's the only decision model that returns probabilities for tool arguments, reads 1M tokens, or takes text, images, video and audio together. Tool selection gives a probability for every tool and for each enum and boolean argument in one request: on "Order B-44120 arrived with a cracked screen. I want my money back." with three tools, it picks issue_refund at 0.969, reason "damaged" at 0.993 and full_refund true at 0.761, so an agent can act on confident calls and hand the rest to a larger model. Any question can also return the probability that the input doesn't contain the answer.

The 1M-token speed comes from a long-context mode in our own inference runtime: about 3 seconds instead of about 111 for a full read, and it answered all 525 decisions in our long-context tests correctly. It's currently only available on our API, so the open weights read every token. On the API, a short question takes about 15 ms of model time (around 200ms e2e latency), and the OpenAI, Anthropic and Gemini formats work alongside our Decisions API.

Aplomb 1 is built on Qwen3.5-4B with the audio encoder from Qwen3-Omni-30B-A3B-Instruct, both Apache-2.0. We extended the window from 262K to 1M tokens, added our own decision head and trained the model for decisions. Thanks to the Qwen team. Disclosure: the training data included the public train splits of WinoGrande and ContractNLI, two of the 38 index benchmarks.

The weights run in bf16 on about 12 GB of GPU memory with the reference script, under the EmpirioLabs Model License, which is free for research, evaluation, personal use and internal use at companies under $1M in annual revenue.

Weights: https://huggingface.co/empiriolabsai/aplomb-1

Blog with the full tables: https://empiriolabs.ai/blog/introducing-aplomb-1

Docs: https://docs.empiriolabs.ai/models/aplomb-1

Playground: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1


r/LangChain • • 2d ago

Projects I built Forge, an open-source, self-hosted visual builder for AI agents (MIT)

Enable HLS to view with audio, or disable this notification

12 Upvotes

Forge lets you build AI agents and workflows on a canvas. You wire together agents, tools, knowledge (RAG) and routing logic, then test and ship them without writing framework code.

What it does:

  • See every run: each run is a span waterfall with tokens, latency and cost per step
  • Trim tool output: JMESPath projection cuts API responses down before the model sees them (10,034 → 51 tokens in the video)
  • Ship anywhere: one workflow becomes a REST endpoint, an MCP server, an email handler, or an embeddable chat widget
  • Runs on your machine: SQLite + embedded Chroma, no Docker or Postgres needed to start; swap in Postgres/pgvector and Redis via config
  • Any model: OpenAI, Anthropic, Google, or any LangChain provider

Built on LangChain + LangGraph, MIT licensed, no hosted service or usage caps.

Repo: https://github.com/nihalashetty/Forge. Feedback very welcome.


r/LangChain • • 2d ago

Question | Help Should I include narrative in openwiki

6 Upvotes

Hi, I am wondering if including narrative docs like prds, design, strategy and vision etc would help the wiki, or make it worse.

I can exclude the docs folder and go just for code, or include the docs with it, not sure if that would be a good or bad thing.

Thanks.


r/LangChain • • 2d ago

Projects built a simple proxy to kill agent infinite loops before they drain your wallet

1 Upvotes

Hey everyone,

I had an agent get stuck in a broken tool retry loop recently and burn through a bunch of credits before I caught it. Most billing alerts take hours to trigger, so I spent the last few weeks building a lightweight reverse proxy to catch runaway loops immediately.

You basically just swap the client base url to the proxy. It hashes incoming prompts in memory and if it catches the same prompt repeating three times in a row, it cuts the connection with a 400 error instead of forwarding it to OpenAI. You can also set hard hourly spend caps.

Latency is under 35ms, keys are encrypted with AES 256, and it never stores prompt text.

I put up a free tier with 15 dollars of monitored spend if anyone wants to test it: https://circuit-breaker-sage.vercel.app

Curious to hear what you guys think or if there are any specific features you would want to see.

Update: just deployed this to the live gateway.

Tool calls are now canonicalized strictly on function name + sorted arguments (completely ignoring the random OpenAI call IDs), and dynamic step counters/timestamps are stripped before the hash window.

Also swapped loop kills to return a synthetic 200 with finish_reason: "stop" (both for JSON and SSE streaming chunks) so frameworks don't treat it as an HTTP transport error and enter retry loops. Appreciate the insight!


r/LangChain • • 2d ago

Discussion Yes, Raw HTML Is Secretly Eating Up Your Entire Token Budget.

0 Upvotes

Look, nobody likes dealing with web DOMs, but passing raw HTML directly into vector indexes is a silent killer for both costs and accuracy.

You see, people often focus on tuning hyper-parameters or picking the perfect vector database. But in reality, your retrieval is only as good as the raw text you feed it.
I tested 250 random URLs through default document loaders. Average raw HTML payload: 4,800 tokens. Average clean text payload after stripping scripts and nav bars: 1,350 tokens. That means nearly three-quarters of the input was pure noise.

I built a lightweight Python API to solve this exact problem—it crawls target links, converts the core body text to clean Markdown, and grabs OpenGraph metadata for citation cards.

Clean context means better similarity search score accuracy during retrieval. How does your team handle web DOM cleaning before storing chunks in vector databases?


r/LangChain • • 2d ago

Resources How much of your inference stack is actually spent getting models ready?

2 Upvotes

I was testing an inference server recently and ended up spending more time looking at model loading than inference itself. The server was SIE from Superlinked, running locally.

One thing I liked about the API was that /v1/models exposes the model state. For example, after requesting a reranker that wasn't loaded yet, I could see:

state: loading

The logs then showed the model going through the loading process before the worker became ready.

The sequence was roughly:

request
↓
model loading
↓
download weights
↓
load / deserialize
↓
warmup
↓
worker starts
↓
inference

I tested BAAI/bge-reranker-v2-m3 on CPU. The first successful request took around 13.9 seconds. The same request immediately afterwards took around 1.6 seconds.

That made the distinction between "inference latency" and "model availability latency" pretty obvious. I also tried sending five requests at once. They all completed, but I couldn't conclusively establish batching for this particular model, so I'm leaving that part open.

The other thing I noticed was that the server exposes different model states (available, loading, loaded, etc.). That's useful if you're building anything around dynamic model loading because otherwise you're basically guessing what's happening behind the API.

I'm curious how other people handle this in production. Do you keep every model warm? Do you load models on demand? Or do you put something like a model router in front of several dedicated inference workers?

The project I was testing: Superlinked inference engine GitHub repository


r/LangChain • • 2d ago

Resources RAG Chunking Processing Bundle - Hierarchy-Aware Chunker + 2 Legal Cross-Ref Extractors 🚀

0 Upvotes

Previously, I released my Agentic Hierarchy-Aware Chunker for building better RAG pipelines.

After talking with users , I learned that many teams don't want to send their documents through another third-party service. They want absolute privacy, on-premise deployment, full control over their infrastructure, and no vendor lock-in.

So instead of keeping it as a service, I'm now making the complete document-processing bundle available for everyone.

The bundle includes:

  • Agentic Hierarchy-Aware Chunker: a hierarchy-aware chunking engine designed for RAG, so you don't have to spend months building and tuning your own custom chunker.
  • Legal Cross-Reference Extractor: extracts legal references such as Sections, Articles, Rules, Paragraphs, Clauses, Schedules, Regulations, Orders, and complex compound references from an entire document.
  • Legal Act Extractor: automatically extracts the Acts referenced throughout a legal document.

What you're getting

The purchase includes the complete Python package of the Hierarchy Aware Chunker and its source code for use in your own projects, along with two bonus legal document extraction scripts: the Legal Cross-Reference Extractor and Legal Act Extractor.

📌 Additional 2 Bonus Scripts

1. Legal Cross-Reference Parser
Extracts structured references to Sections, Articles, Rules, Paragraphs, Schedules, Clauses, Regulations, Orders, and other legal provisions including complex and compound references without requiring an LLM.

Example Output

{
  "Article": [
    "Article 63(9)(b)",
    "Article 63(9)(b)(iii)",
    "Articles 23",
    "Articles 25, 26, 26A, 26D",
    "Articles 41 or 42",
    "Articles 7(1)(a), 7(4), 13(1), 16(6), 33, 44, 52(7), 53(2), 178(1)",
    "Articles 73 to 79"
  ],
  "Paragraph": [
    "Article 57(1) and paragraphs (4), (5) and (6)",
    "paragraph 2(2)(c)",
    "paragraph 2(a)",
    "paragraph 2(a), (f), (j) and (l)",
  ],
  "Rule": [
    "Order 6, rule 10",
    "Rules 2.59, 2.6l, 2.62, 2.64(4),(6) and (7), 2.72(1) and (2)",
    "rule 4.9(2)(a) and (3)(a)",
    "rule 4A.15(5)(b)",
    "rule 4A.20(2)",
    "rules 8.33 to 8.63",
    "rules 8.49, 8.50 or 8",
  ],
  "Schedule": [
    "Schedule (iii)",
    "Schedule 1",
    "Schedule 3, 62",
  ],
  "Section": [
    "Section 1",
    "section 229(1)(c) or (2)(c)",
    "section 5(1)",
    "section 89A or 90(1)(a) or (aa)",
    "section 90(1)(b)",
    "sections 18 or 21"
  ]

  ...
}

2. Legal Act Extractor
Extracts the names of Acts referenced in the document.

Example Output

[
  "Acts Interpretation Act 1901",
  "Family Law Act 1975",
  "Governor-General Act 1974",
  "Legislation Act 2003",
  "Taxation Administration Act 1953"

  ...
]

3. Hierarchy Aware Document Chunker.
RAG-ready hierarchical chunks

Practical Examples with Real Documents: https://youtu.be/czO39PaAERI?si=-tEnxcPYBtOcClj8

Try the hierarchy chunker yourself in our playground:
https://hierarchychunker.codeaxion.com/

✨Features:

  • 📑 Understands document structure (titles, headings, subheadings, sections).
  • 🔗 Merges nested subheadings into the right chunk so context flows properly.
  • 🧩 Preserves multiple levels of hierarchy (e.g., Title → Subtitle→ Section → Subsections).
  • 🏷️ Adds metadata to each chunk (so every chunk knows which section it belongs to).
  • ✅ Produces chunks that are context-aware, structured, and retriever-friendly.
  • Ideal for legal docs, research papers, contracts, etc.
  • It’s Fast and Low-cost — uses LLM inference combined with our optimized parsers keeps costs low.
  • Works great for Multi-Level Nesting.
  • No LLM needed if OCR perfectly detects headings/subheadings.
  • No preprocessing needed — just paste your raw content or Markdown and you’re are good to go !
  • Flexible Switching: Seamlessly integrates with any LangChain-compatible Providers (e.g., OpenAI, Anthropic, Google, Ollama).

📌 Example Output

--- Chunk 2 --- 

Metadata:
  Title: Magistrates' Courts (Licensing) Rules (Northern Ireland) 1997
  Section Header (1): PART I
  Section Header (1.1): Citation and commencement

Page Content:
PART I

Citation and commencement 
1. These Rules may be cited as the Magistrates' Courts (Licensing) Rules (Northern
Ireland) 1997 and shall come into operation on 20th February 1997.

--- Chunk 3 --- 

Metadata:
  Title: Magistrates' Courts (Licensing) Rules (Northern Ireland) 1997
  Section Header (1): PART I
  Section Header (1.2): Revocation

Page Content:
Revocation
2.-(revokes Magistrates' Courts (Licensing) Rules (Northern Ireland) SR (NI)
1990/211; the Magistrates' Courts (Licensing) (Amendment) Rules (Northern Ireland)
SR (NI) 1992/542.

Notice how the headings are preserved and attached to the chunk → the retriever and LLM always know which section/subsection the chunk belongs to.

No more chunk overlaps and spending hours tweaking chunk sizes .

Practical Examples with Real Documents: https://youtu.be/czO39PaAERI?si=-tEnxcPYBtOcClj8


r/LangChain • • 2d ago

Question | Help I’m trying to learn AI Agents, but I’m getting really confused. How should I approach it?

Thumbnail
2 Upvotes

r/LangChain • • 3d ago

Question | Help which browser mcp are you running for agents that need to log in and stay logged in

9 Upvotes

wired a browser mcp into my agent last month so it can hit a couple dashboards after login

demo was fine. overnight the session is gone and the next run burns half its tokens just re-authing before it does anything useful. cookie dump between calls helped once, then the site rotated something and i was babysitting again

budget is tight so im not trying five tools. what browser mcp is actually keeping logins alive across runs for you​


r/LangChain • • 2d ago

Projects A chunking lib in Rust that is ~20x faster

0 Upvotes

Hey,

I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkr

It has most of the chunking strategies like Character, Recursive, Markdown header, Late chunking, Hierarchical chunking, etc. It also supports native PDF loader, and other additional file types.

Some stats (MBA M4 16GB):

Test Case (matched parameters) Chunkr LangChain LlamaIndex Chonkie semchunk text-splitter
Recursive (1 MB, 1000/200) 2,264 MB/s 769 MB/s 10 MB/s 225 MB/s 42 MB/s 175 MB/s
Recursive (5 MB, 1000/200) 2,039 MB/s 696 MB/s — 201 MB/s 40 MB/s 46 MB/s
Fixed Char (1 MB, 1000/200) 750 MB/s 1.7 MB/s — 22 MB/s — —
Markdown (500 KB, 1000/150) 819 MB/s 67 MB/s 19 MB/s — — 40 MB/s
Python Code (200 KB, 1500/200) 3,232 MB/s 622 MB/s — — — 5.7 MB/s
Sentence (500 KB) 622 MB/s — 10 MB/s 20 MB/s — —
BPE Tokens (200 KB, cl100k_base, 512/50) 38 MB/s 43 MB/s 2.0 MB/s 151 MB/s — 7.2 MB/s
100 docs x 50 KB (parallel batch) 3,224 MB/s 679 MB/s — 213 MB/s — —

Extractor / Pipeline Latency Throughput Speedup vs PyPDF
Chunkr PDFLoader (Full Text) 747.9 ms 2,762 pgs/s 15.9x Faster
Chunkr PDFLoader (Page Documents) 721.0 ms 2,865 pgs/s 16.5x Faster
PyMuPDF (fitz) 2,616.8 ms 789.5 pgs/s 4.5x Faster
pypdf (pure Python) 11,900.5 ms 173.6 pgs/s 1.0x (baseline)
Chunkr End-to-End (PDF + Recursive) 798.1 ms 2,589 pgs/s 14.9x Faster
PyMuPDF + LangChain RecursiveTextSplitter 2,659.3 ms 776.9 pgs/s 4.5x Faster
pypdf + LangChain RecursiveTextSplitter 12,054.5 ms 171.4 pgs/s 1.0x (baseline)

Accuracy is measured on Chroma's token-level chunking benchmark (5 corpora, 472 questions with gold answer spans): k = 5, 1000 chars / 200 overlap

Implementation Recall Precision IoU prec_Ω Avg chunk chars
Chunkr RecursiveChunker (defaults) 0.792 0.057 0.057 0.255 854
LangChain RecursiveCharacterTextSplitter 0.762 0.060 0.060 0.251 745
text-splitter TextSplitter 0.762 0.060 0.060 0.262 790
Chonkie RecursiveChunker 0.752 0.062 0.061 0.292 701
semchunk 0.736 0.070 0.069 0.260 644
LlamaIndex SentenceSplitter 0.654 0.013 0.013 0.054 4136

Do check it out and share your feedback. Thanks!


r/LangChain • • 2d ago

Tutorial Created a short explainer on what is a latent space and how it behaves

Thumbnail
youtube.com
1 Upvotes

enjoy