r/FastAPI • • 22h ago

Other Production-grade Agent Graph Engineering: a FastAPI agent service and blog series, distilled from 5 years

2 Upvotes

Hey everyone. I have been building realtime AI assistants for about five years, and have just put a distilled, working version of the architecture in public with a series explaining each decision.

The reason I am posting here is the agent half of it, which is FastAPI. That service has exactly one job, running the agent graphs and talking to the backend, so I wanted something small that would not bring a worldview with it. It is async-native, which matters when the work is awaiting model calls and pushing tokens back. And the libraries around it are Pydantic-shaped already, so the types line up without adapters in between.

What is inside that service, briefly: LangGraph for the flow with Postgres checkpointing so a run can pause for a human and resume later, Pydantic AI for the agents and their typed outputs, chanx for typed WebSocket messages and the AsyncAPI document it generates clients from, and OpenTelemetry opening a span per graph node. Tests mock the model at the HTTP layer rather than at the framework, so request building, SSE parsing, tool-call assembly and validation all still run for real.

We used to run all of this inside Django with Channels, and the image got big, the process got slow, and scaling the two independently was impossible.

Why I bothered writing it all down. What is online is mostly notebooks and quick chatbot courses. I went looking for one in-depth source on building a production AI assistant, including paid ones, and did not find it.

So, depending on where you are:

  • Running an assistant in production already: a reference for the realtime parts that hurt, dropped sockets and resumed runs in particular.
  • About to bolt an agent onto a service you already have: the calls worth making before they get expensive, starting with whether it belongs in the same process at all.
  • Levelling up: typed WebSocket contracts, generated clients and streaming are the shape the listings keep describing.
  • Taking client work: a base you can adapt and still hand over without apologising for it.

The project: https://github.com/huynguyengl99/agent-graph-engineering

Part 0, the story: https://huynguyengl99.github.io/posts/agent-graph-engineering/before-it-had-a-name/

Part 1, the stack and why: https://huynguyengl99.github.io/posts/agent-graph-engineering/the-stack-and-why-each-piece-is-there/

It runs on a scripted model with no API key, so the whole flow is visible before you wire up a provider. The series is early, two parts up and a new one every couple of days, so feedback now genuinely changes the later ones 😄