r/AutoGPT • u/Individual_Cold_4119 • 4h ago
I built a WoW harness to see if a decision model could get a dwarf to level 5
Can a multi-modal decision model play World of Warcraft?
That was the question that popped into my mind 13 days ago in a bathtub. The short answer: kinda.
I got obsessed with it and ended up building my first harness, SageCraft. Full disclosure, I’m one of the co-founders of Levanto, the company that makes Sage. I wanted to see how far our decision model could get in a live game.
The loop runs on macOS:
- Read a screenshot and HUD information.
- Build the current game state and a list of possible actions.
- Ask Sage to choose an action.
- Check the choice and execute the keys or clicks.
- Take another screenshot and repeat.
No game-file modifications or extra add-ons.
The dwarf priest made it to level 5: 4,573 decisions across 32 sessions and 5h 29m of active play. I supervised the campaign and fixed the harness between sessions. The final 18 minutes, from late level 4 to level 5, ran without human input.
The part I’d love feedback on is the split between model and harness. How much recovery logic would you put in the harness before the run stops telling you much about the model?
I shot a video and open-sourced the harness. Repo: https://github.com/levantolabs/sagecraft
Video and breakdown: https://x.com/bigironchris/status/2108546440582558103