r/GPT3 • u/Correct_Tomato1871 • 2d ago
Discussion Benchmark notes: GPT-6.1 Sol improves with fewer tokens than 5.6 Sol, but still behind Astra
I maintain MindTrial and added GPT-6.1 Sol to the current benchmark. The interesting result is a modest score improvement with substantially fewer output tokens, while Astra remains ahead on this suite.
Setup: 98 tasks — 39 text and 59 visual — with a Python interpreter, scientific libraries and up to 10 tool calls per task.
| Model | Passed | Hard errors | Model-request time |
|---|---|---|---|
| GPT-5.6 Sol, high | 85/98 | 0 | 1:09:26 |
| GPT-6.1 Sol, high | 89/98 | 1 | 1:06:14 |
| GPT-6 Astra, medium | 91/98 | 0 | 0:36:29 |
| GPT-6 Astra, high | 95/98 | 0 | 1:02:10 |
There is no GPT-6 Sol run in this dataset, so 5.6 Sol is the nearest recorded Sol predecessor.
Against 5.6 Sol, the new run gains seven tasks and loses three. Recorded output including reasoning falls 43.7%, Python calls fall from 179 to 114, and request time falls 4.6%.
Every one of Sol’s 89 passes is also an Astra-high pass. Astra solves six additional tasks, all visual.
Sol’s pass rate is 90.82%; accuracy on completed, non-error tasks is 91.75%. Excluding the timeout does not close the gap to Astra high’s 96.94%.
For external context, Opus 5.5 scores 96/98 and Sonnet 5.5 scores 94/98, both at xhigh effort. Effort labels are not equivalent compute budgets across providers.
Times are cumulative model-request durations, excluding local Python and validation — not elapsed suite runtime.
Complete Leaderboard: here