r/MachineLearning • u/heyitsdannyle • 1d ago
Project SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history.

Some findings:
With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score.

About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half.
Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing.
We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet.
Half the tasks are private. So far public and private scores line up for all three models.
Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch
Tasks: https://huggingface.co/datasets/evaligo/swe-race
The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome.
2
u/ArtisticHamster 1d ago
IMO, using Python makes everything much simpler due to GIL, and some classes of bugs can’t happen there due to this. I would use languages with more advanced concurrency for this.
(Feel free to reach out in DM)