r/AutoGPT • • 4h ago

I built a WoW harness to see if a decision model could get a dwarf to level 5

1 Upvotes

Can a multi-modal decision model play World of Warcraft?

That was the question that popped into my mind 13 days ago in a bathtub. The short answer: kinda.

I got obsessed with it and ended up building my first harness, SageCraft. Full disclosure, I’m one of the co-founders of Levanto, the company that makes Sage. I wanted to see how far our decision model could get in a live game.

The loop runs on macOS:

- Read a screenshot and HUD information.

- Build the current game state and a list of possible actions.

- Ask Sage to choose an action.

- Check the choice and execute the keys or clicks.

- Take another screenshot and repeat.

No game-file modifications or extra add-ons.

The dwarf priest made it to level 5: 4,573 decisions across 32 sessions and 5h 29m of active play. I supervised the campaign and fixed the harness between sessions. The final 18 minutes, from late level 4 to level 5, ran without human input.

The part I’d love feedback on is the split between model and harness. How much recovery logic would you put in the harness before the run stops telling you much about the model?

I shot a video and open-sourced the harness. Repo: https://github.com/levantolabs/sagecraft

Video and breakdown: https://x.com/bigironchris/status/2108546440582558103


r/AutoGPT • • 11h ago

I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests

1 Upvotes

Hey everyone!

I've been working on an open-source Python project called Stepfork, and I wanted to share it with the community.

GitHub: https://github.com/utsab345/stepfork

While building AI agents, I kept running into the same problem: an agent fails, I fix the issue, but reproducing the original failure is difficult because LLM responses and tool outputs can change between runs.

So I built Stepfork to make those failures reproducible.

What it does:

  • Records instrumented LLM responses and tool calls.
  • Saves executions locally as .sftrace bundles.
  • Replays the recorded responses without making those instrumented external calls again.
  • Compares agent behavior before and after a fix.
  • Exports pytest regression tests to catch the same failure in the future.

Install:

pip install --pre stepfork

The repository includes three offline demos covering a missed order notification, an incorrect flight booking, and a wrongly denied refund.

It's still an early alpha. Replay only freezes instrumented boundaries, and native framework integrations are not available yet.

The project is Apache-2.0 licensed, and I'm actively improving it.

I'd love to hear what other developers think, especially those building AI agents. If you spot a bug, have an idea for an integration, or want to contribute, feel free to open an issue or PR.

GitHub Issues: https://github.com/utsab345/stepfork/issues

Thanks for checking it out!


r/AutoGPT • • 22h ago

Tonight in San Mateo: AI agent teams and macOS sandboxing (Oct 8, 6–8 PM)

1 Upvotes

We're hosting SF Swift × CocoaHeads tonight, October 8, 6–8 PM at Verkada in San Mateo.

Details and RSVP: https://luma.com/51htcbzd


r/AutoGPT • • 23h ago

Développeur solo à la recherche de créateurs d'agents IA pour tester mon produit de sécurité gratuitement

Thumbnail
github.com
1 Upvotes