SWE-bench Verified
Real software-repair tasks
No result yetA continuously improving model
Echo is not perfect yet. Its early results are already impressive: it matches Claude Fable on several evaluations, while other results show where we still have work to do. We publish both.
Echo keeps improving. We continuously test new open-weight models, improve how they work together, and expand this evaluation set. This page is updated as new results are ready, while older results remain visible.
Latest measured results
These are our newest score summaries. Question-by-question examples are still being added, so each card says exactly what is and is not available.
Echo evidence
Every MATH-500 result is inspectable. This view contains the exact 500 questions, stored Echo and Fable answers, and recorded grades behind the headline scores.
| # | Question | Correct answer | Echo | Fable |
|---|
Open every result
Every card below lets you inspect the exact questions and stored answers. A direct comparison appears only when Echo and Fable answered the same questions.
These scores come from the recorded runs shown here. Each row preserves one run per available system. A new run of the same prompt can produce a different answer.
Benchmarks are ordered by performance difference.
Next evaluations
These cards are placeholders, not claims. Results will appear only after the complete evaluation is ready to publish.
Real software-repair tasks
No result yetAbstract problem solving
No result yetPractical coding tasks
No result yetQuestion-by-question results
| # | Question | Correct answer | Echo | Fable reference |
|---|
Transparency
These are our own published evaluations, not an independent third-party review. We show the limits so the numbers are not asked to prove more than they can.
For the detailed results, you can open every question, the correct answer, and both stored model answers. If an answer was not returned, the row says so explicitly and leaves it out of direct comparisons. The page is read-only and does not rerun either model.
Yes. Model outputs can vary between runs. This page preserves the exact measured run used for each displayed score, and every question view shows the number of stored runs.
Only when Echo and Fable were tested on the same questions. If one side is missing answers, the page says the scores are not directly comparable.
No. They show how the tested systems performed on the questions listed here. Public benchmarks can also appear in model training data, so these results are evidence, not a universal guarantee.
Yes. We show recorded losses as clearly as the results where Echo matches or leads. Evaluations still being completed are withheld until their runs are ready.
The 100-question set was used while building Echo, so it is useful but not an independent test. A separate 41-question set was tested later, but Fable answered only 19 of those questions, so the two overall scores cannot be compared directly. The 19 questions both answered remain visible as an exploratory view.
Math and multiple-choice answers are checked against the known answer. Code is run against tests. We keep the checked answers fixed so the displayed totals do not change silently.
The newest score summaries passed our arithmetic checks, but their question-by-question answers are not yet published here. They remain clearly labeled as summaries until those examples are added.
For developers
Read-only access to the same results shown on this page.
curl {base}api/current
curl {base}api/current/MATH-500/evidence
curl '{base}api/current/MATH-500/rows?offset=0&limit=50&outcome=all'
curl {base}api/current/MATH-500/rows/math500:000:13bb5d32a5afacb3
curl {base}api/benchmarks
curl '{base}api/benchmarks/MATH-500/rows?offset=0&limit=50'
curl {base}api/rows/math500:0280d332515ce2d26069
curl {base}api/manifest