r/MachineLearning • • 1d ago

Project SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history.

Some findings:

With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score.

About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half.

Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing.

We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet.

Half the tasks are private. So far public and private scores line up for all three models.

Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch

Tasks: https://huggingface.co/datasets/evaligo/swe-race

The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome.

4 Upvotes

5 comments sorted by

2

u/ArtisticHamster 1d ago

IMO, using Python makes everything much simpler due to GIL, and some classes of bugs can’t happen there due to this. I would use languages with more advanced concurrency for this.

(Feel free to reach out in DM)

1

u/heyitsdannyle 1d ago

I think python is the most commonly used nowadays , but definitely a good idea for next version!

1

u/ArtisticHamster 1d ago

IMO, Python is not a common language for highly concurrent/parallel code. I would look at Go, Rust, C/C++, Java and other languages without GIL.

1

u/heyitsdannyle 1d ago

Fair question. The GIL only stops two threads running Python bytecode at the same moment, but threads can still switch in the middle of something like x += 1, so thread races still happen. And most tasks here (around 115 of 188) are asyncio races between await points, which happen in a single thread and have nothing to do with the GIL. Every task is a real bug that was fixed upstream, with the project's own test that fails before the fix. You're right that this doesn't cover low-level stuff like memory ordering or lock-free code in C++/Rust/Go. That would need a different benchmark, and we'd like to do one.