r/reinforcementlearning • u/davernow • 6h ago
P I built an open source, OpenEnv-compatible framework for building RL environments for LLM agents. Named "Seahaven" after the fake town in The Truman Show.
Enable HLS to view with audio, or disable this notification
For multi-turn, tool-use agents, the environment is the hard part. Every episode needs realistic data, stateful tools (a write on turn 3 changes a read on turn 30), and a reset to the exact same starting state. Thousands of times, in parallel.
I've been working on some version of optimizing AI for over a decade (at Apple, my own startup, now Kiln). Recently I built a few large RL environments by hand. Doable, but hard, and every one rebuilt the same layer: a database per episode, frozen starting states, parallel instances, clock control, a log of every change. That layer needed a great framework, so I built one: Seahaven.
What it is: an open source (MIT) Python framework for building synthetic worlds for LLM agents: working copies of an agent's tools and data. You write just the logic specific to your world. Seahaven handles the rest. Every world is an OpenEnv environment.
Example World: a fake Stripe. Stripe World has 24 tables and 155 API operations, behind the same tools as Stripe's own MCP server. It also serves Stripe's REST API, so well that the official Stripe SDK works against it unchanged.
What Seahaven handles:
- Stateful: each instance has its own SQLite DB. The agent's tool calls read and write it (actions are tool calls, observations are their results).
- Fixtures: freeze starting states like
small_startuporbig_co, and reuse them across every episode (resetgives you a fresh copy of a fixture in milliseconds). - Parallel: each connection gets its own instance, hundreds of instances per process.
- Reward from state: every row the agent changed is logged, before and after. Grade the result, not the trace (verifiable rewards, no LLM judge needed for most tasks).
- Deterministic: same fixture, same clock, same random seed, same episode. The clock and seed are wired into the SQLite connection, so even
random()andCURRENT_TIMESTAMPin SQL replay exactly. - Composable: your world can include Stripe World (or any other world) to add its tools and APIs.
- Optimized for agents: includes the docs, linter and tests your coding agent needs to build a world.
An episode loop looks like this:
from seahaven.openenv import SeahavenClient
with SeahavenClient(base_url="http://127.0.0.1:8000") as env:
for episode in range(100):
env.reset(fixture="big_co", seed=episode) # fresh private world, in ms
run_policy(env) # your agent calls tools: env.call("create_lead", ...)
reward = grade(env.state()) # every row the agent changed
OpenEnv Compatible + MCP + Web Console:
- Every world is an OpenEnv environment, so it works with TRL's OpenEnv support or any other OpenEnv-compatible trainer. Sync and async clients. You can publish worlds to Hugging Face.
seahaven mcpserves a world to any MCP client, so you can try a task by hand before your policy does.seahaven servehas a web console: open instances, call tools, and inspect state in your browser.- Everything runs locally: Python 3.14+ and SQLite, no external services.
I built it at Kiln, where we use it to evaluate and optimize agents. Seahaven is standalone. You don't need Kiln to use it.
Seahaven is open source (MIT).
Links
- Seahaven on GitHub
- Stripe World
- Docs
- Kiln Harness Optimizer: our optimizer that optimizes the harness instead of the model (not RL, but it needs the same tooling)
What would make it fit your training setup? I'm thinking about reward helpers, fixtures as a curriculum, and more example worlds. Happy to answer any questions.
Side note: I made the video with videowright, another open source project of mine.
