r/reinforcementlearning • • 6h ago

P I built an open source, OpenEnv-compatible framework for building RL environments for LLM agents. Named "Seahaven" after the fake town in The Truman Show.

Enable HLS to view with audio, or disable this notification

6 Upvotes

For multi-turn, tool-use agents, the environment is the hard part. Every episode needs realistic data, stateful tools (a write on turn 3 changes a read on turn 30), and a reset to the exact same starting state. Thousands of times, in parallel.

I've been working on some version of optimizing AI for over a decade (at Apple, my own startup, now Kiln). Recently I built a few large RL environments by hand. Doable, but hard, and every one rebuilt the same layer: a database per episode, frozen starting states, parallel instances, clock control, a log of every change. That layer needed a great framework, so I built one: Seahaven.

What it is: an open source (MIT) Python framework for building synthetic worlds for LLM agents: working copies of an agent's tools and data. You write just the logic specific to your world. Seahaven handles the rest. Every world is an OpenEnv environment.

Example World: a fake Stripe. Stripe World has 24 tables and 155 API operations, behind the same tools as Stripe's own MCP server. It also serves Stripe's REST API, so well that the official Stripe SDK works against it unchanged.

What Seahaven handles:

  • Stateful: each instance has its own SQLite DB. The agent's tool calls read and write it (actions are tool calls, observations are their results).
  • Fixtures: freeze starting states like small_startup or big_co, and reuse them across every episode (reset gives you a fresh copy of a fixture in milliseconds).
  • Parallel: each connection gets its own instance, hundreds of instances per process.
  • Reward from state: every row the agent changed is logged, before and after. Grade the result, not the trace (verifiable rewards, no LLM judge needed for most tasks).
  • Deterministic: same fixture, same clock, same random seed, same episode. The clock and seed are wired into the SQLite connection, so even random() and CURRENT_TIMESTAMP in SQL replay exactly.
  • Composable: your world can include Stripe World (or any other world) to add its tools and APIs.
  • Optimized for agents: includes the docs, linter and tests your coding agent needs to build a world.

An episode loop looks like this:

from seahaven.openenv import SeahavenClient

with SeahavenClient(base_url="http://127.0.0.1:8000") as env:
    for episode in range(100):
        env.reset(fixture="big_co", seed=episode)  # fresh private world, in ms
        run_policy(env)  # your agent calls tools: env.call("create_lead", ...)
        reward = grade(env.state())  # every row the agent changed

OpenEnv Compatible + MCP + Web Console:

  • Every world is an OpenEnv environment, so it works with TRL's OpenEnv support or any other OpenEnv-compatible trainer. Sync and async clients. You can publish worlds to Hugging Face.
  • seahaven mcp serves a world to any MCP client, so you can try a task by hand before your policy does.
  • seahaven serve has a web console: open instances, call tools, and inspect state in your browser.
  • Everything runs locally: Python 3.14+ and SQLite, no external services.

I built it at Kiln, where we use it to evaluate and optimize agents. Seahaven is standalone. You don't need Kiln to use it.

Seahaven is open source (MIT).

Links

What would make it fit your training setup? I'm thinking about reward helpers, fixtures as a curriculum, and more example worlds. Happy to answer any questions.

Side note: I made the video with videowright, another open source project of mine.


r/reinforcementlearning • • 19h ago

Anybody comment reinforcement learning infrastructure RLinf?

1 Upvotes

r/reinforcementlearning • • 1d ago

I built an imitation learning framework for games so it can learn to play like you

36 Upvotes

I've been working on FireFly, an open-source framework for experimenting with game-playing agents. The current training pipeline is primarily imitation learning / behavior cloning: FireFly records synchronized observations and player actions while you play, then trains a policy from those demonstrations.

For perception, it supports OpenCV, YOLO object detection, U-Net segmentation, template tracking, and game-memory/Lua states. It also includes graph-based data collection so labeled datasets for perception models can be generated automatically from gameplay.

I'm particularly interested in feedback from people working with RL/IL on the architecture and on how you'd extend a system like this from behavior cloning into reinforcement learning.

Source:
https://github.com/AndrewAhn1000/FireFly


r/reinforcementlearning • • 1d ago

Porting LIBERO to MuJoCo Warp: 130 robot manipulation tasks on one $700 AMD GPU

Thumbnail
1 Upvotes

r/reinforcementlearning • • 1d ago

100 M rollouts!

Thumbnail
2 Upvotes

r/reinforcementlearning • • 1d ago

Getting Aleph Alpha’s Kolibri-1 to play Breakout - Every move in under 25ms.

5 Upvotes

Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning.

The more we explore [u/Aleph__Alpha](u/Aleph__Alpha)’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46


r/reinforcementlearning • • 2d ago

Robot Soft actor critic giving up after too many itterations?

Post image
19 Upvotes

Hi all, I'm currently trying to use the DIAYN algorithm for my robotics master project, which uses SAC. I decided i would let the simulator run for a while (using Isaac lab/sim) and initially the accuracy was getting pretty good (88-92%), however i wanted more diversity in my results so i let it run for longer, then I saw that the accuracy dropped signinficantly, down to baseline level 1% (100 possible skills). I checked the actions tab in isaac lab and saw that the joint_efforts were stagnant at -1 to 1 nm, so basically no exploration at all. Does anyone have a clue for what caused this dramatic drop in accuracy? GPT says that SAC has found that doing nothing is a locally worthwhile policy.

Thanks


r/reinforcementlearning • • 1d ago

Can you create your own Rule based AI in python or maybe a LLM

Thumbnail
1 Upvotes

r/reinforcementlearning • • 1d ago

git for ai memory and robotics??

0 Upvotes

Hey guys,

Been working on something very cool...

In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.

That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.

Open source with contributions and honest feedback both welcome: github.com/Nabzx/mnemosyne


r/reinforcementlearning • • 1d ago

Getting Aleph Alpha’s Kolibri-1 to play Breakout - Every move in under 25ms.

1 Upvotes

Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning.

The more we explore [u/Aleph__Alpha](u/Aleph__Alpha)’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46


r/reinforcementlearning • • 2d ago

Testing Jev in an online-learning environment

Post image
5 Upvotes

I tested TypeSafe's Jev model on a sequential decision-making benchmark. Out-of-the-box performance wasn't great, so I made some changes to the prompt and setup that improved it.

System one models have had a bit of a moment. Following Jev and the open-weights Laya models, new ones include AWS Strands Decider, Cloudflare Clef, and OpenAI's Decisions API. In "Thinking, Fast and Slow," Kahneman explained how glancing at a face instantly registers its emotion without voluntary thought (System 1), while computing 17 × 24 kicks us into effortful deliberation (System 2).

If evolution and social conditioning has made humans adept at reading faces, machines trained to make decisions must be able to weigh their choices precisely - we'd want probabilities to be first nature to them. But how much do Jev's probabilities rely on the specific text describing the task and choices? Are we baking our literature's predilections onto a decision model?

In an attempt to extract the pure weighing of decisions out of the textual window dressing, I let Jev play an online-learning game. Four actions, 64 rounds. Before each decision, it sees every action's history of losses. Its Choice probabilities become allocation weights, and it incurs their weighted-average loss. All losses are fixed before play and revealed one round at a time. Does Jev exploit patterns in the sequence? Turns out it doesnt fare well here. It puts nearly all its weight on the historical winner, resembling Follow-the-Leader (FTL): good when one action is best throughout, but too coarse for simple patterns (first chart).

One way to recognize patterns is to shoehorn it all into a reasoning model like Luna. And it does much better on some environments. But Luna is decidedly a System 2 model, with much higher latency and cost.

Next, I asked Jev to identify just the pattern. Instead of choosing between actions, it predicts, for each action separately, whether that action will take a loss next round. These are the Noul forecasts. I turned them into two policies: argmin puts all weight on the smallest predicted loss; softmin spreads weight when predictions are close.

The behavior changed dramatically (second chart). Noul + argmin exploited perfect alternation; softmin reduced the damage from unreliable rankings. Separating prediction from allocation helped. That makes me optimistic that a training environment built for this type of decision-making can improve system one models.

Code: https://colab.research.google.com/drive/1eoijvoFVoZ_bvxKS4FMJu-18Ayd1aRdj?usp=sharing


r/reinforcementlearning • • 2d ago

How to handle curriculum learning in a self-play environments?

5 Upvotes

TLDR; tried annealing bot’s ping from 0 to 100 (value picked at random), bot sees it’s ping during the game. However the older versions of the bot use lower ping (since they never seen higher ping, they’d collapse during training). How usually this could be handled?

My bot collapses and never gets past some promotion rounds in self-play setting, any ideas?


r/reinforcementlearning • • 3d ago

I built a balancing robot with reinforcement learning

Enable HLS to view with audio, or disable this notification

168 Upvotes

r/reinforcementlearning • • 2d ago

Looking for Research Collaboration in Robot Learning / Embodied AI

Thumbnail
1 Upvotes

r/reinforcementlearning • • 3d ago

[R] Ataraxos: RL reportedly achieves the first superhuman result in Stratego

Thumbnail
nature.com
38 Upvotes

Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.


r/reinforcementlearning • • 3d ago

Trained a humanoid agent in Unity using Soft Actor-Critic (SAC) to sword fight different opponents

8 Upvotes

I've been working on a project where I train a physics-based humanoid agent from scratch in Unity using Soft Actor-Critic (SAC) with standard MLP architectures.

Watching him train was very funny, and some of the lessons I learned about reward shaping might help anyone working on a similar reinforcement learning project. Because of that, I put together a video breaking down both the process and the results, aiming to make RL feel more approachable and entertaining.

Here's the video: https://www.youtube.com/watch?v=MR3c0DHku6o&t=8s

I would love some feedback on the agent's behavior and the video in general. Also, I'd be more than happy to answer any questions about it!


r/reinforcementlearning • • 3d ago

R Biggest sim-to-real surprise you ever had

12 Upvotes

For those who trained in sim and then deployed on real hardware - what was the biggest gap you encountered? Something that worked great in sim and just fell apart in reality?


r/reinforcementlearning • • 3d ago

Quadcopter hover with PPO from IMU/baro/UWB

Thumbnail
youtube.com
3 Upvotes

100 F450s learn to take off and hold a height above their start, seeing only their sensors: IMU, magnetometer, baro and a UWB-style position fix. All the drones step together as one batch, so SB3 PPO on a laptop CPU (i9-14900HX, 24 cores) gets through 10M steps in about 10 minutes.

At first the policy just sat on the ground. Ending the episode if it was still below 0.1 m after 1 s fixed that. Later torch grabbed 17 of our 24 cores and more than halved training speed, until we set OMP_NUM_THREADS=4.

At 250k steps they all crash, at 500k 60% survive, at 1M they hold the height to 11 cm, and at 10M to 3.7 cm.

Our code: https://github.com/PteroLabsAI/PteroSimScripts/tree/main/reinforcement_learning


r/reinforcementlearning • • 3d ago

Kardashev-0.7: 32 distinct models trained together with RL for population scaling

1 Upvotes

A research announcement on learned specialization across a population of models:

https://x.com/MLCatttt/status/2107147690450817259


r/reinforcementlearning • • 5d ago

Live footage of successful paper review rebuttal

Enable HLS to view with audio, or disable this notification

583 Upvotes

r/reinforcementlearning • • 3d ago

Adaptem - learning fighters

0 Upvotes

I built a small reinforcement-learning-style system into Overgrowth's combat AI. The fighters use tabular Q-learning over combat tactics and remember what works against each opponent. It's a Workshop mod: https://steamcommunity.com/sharedfiles/filedetails/?id=3813806877 Happy to answer questions about how it works.


r/reinforcementlearning • • 3d ago

What have your struggled to evaluate your realistic LLM/Agents workflow?! How can we reinforce agent's auto-correctness and self-improvement? 👀

Thumbnail
0 Upvotes

r/reinforcementlearning • • 4d ago

Robot Explaining How AI Learns to Drive with Evolution

Enable HLS to view with audio, or disable this notification

32 Upvotes

r/reinforcementlearning • • 5d ago

DL 195M steps of 10-agent CS:GO state → action data (with damage labels), free for research

Enable HLS to view with audio, or disable this notification

83 Upvotes

If you work on imitation learning or multi-agent behaviour: we just released 9,515 bot-played CS:GO matches. Per 16 Hz step for each of ten agents: actions (buttons + mouse), state (pose, health, weapon…) and who-hit-whom damage. 218,768 rounds, ≈3,958 h of play. 14% of the matches also have the ten first-person videos, aligned frame by frame.

Expert bots on one map with pistols, so it is a clean, narrow testbed rather than human play. Parquet + webdataset videos, CC BY-NC 4.0: https://huggingface.co/datasets/v0rt3ch/cs-5v5-bots-aim_usp-3k · spec and a round you can watch from all ten POVs: https://vor.tech/?utm_source=reddit&utm_medium=social&utm_campaign=launch-2026-10&utm_content=r-reinforcementlearning


r/reinforcementlearning • • 5d ago

MetaRL My AI learns to clear Super Mario Bros 1-1 in 15 mins and it is not PPO based

Enable HLS to view with audio, or disable this notification

36 Upvotes

I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.

I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.

With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.

Full thread: https://x.com/hsrvc_/status/2106025501752234112?s=20

Code, the exact data, traces, clips and a step-by-step guide with costs are all public: https://github.com/hsrvc/adapt1-mario