SWE-bench Verified
Real software-repair tasks
No result yetA continuously improving model
Echo is not perfect yet. Its early results are already impressive: it matches Claude Fable on several evaluations, while other results show where we still have work to do. We publish both.
Echo keeps improving. We continuously test new open-weight models, improve how they work together, and expand this evaluation set. This page is updated as new results are ready, while older results remain visible.
Latest measured results
These are our newest score summaries. Question-by-question examples are still being added, so each card says exactly what is and is not available.
Open every result
Every card below lets you inspect the exact questions and answers. A direct comparison appears only when Echo and Fable answered the same questions.
Benchmarks are ordered by performance difference.
Next evaluations
These cards are placeholders, not claims. Results will appear only after the complete evaluation is ready to publish.
Real software-repair tasks
No result yetAbstract problem solving
No result yetPractical coding tasks
No result yetQuestion-by-question results
| # | Question | Correct answer | Echo | Fable reference |
|---|
Transparency
These are our own published evaluations, not an independent third-party review. We show the limits so the numbers are not asked to prove more than they can.
For the detailed results, you can open every question, the correct answer, and both model answers. The page is read-only and does not rerun either model.
Only when Echo and Fable were tested on the same questions. If one side is missing answers, the page says the scores are not directly comparable.
No. They show how the tested systems performed on the questions listed here. Public benchmarks can also appear in model training data, so these results are evidence, not a universal guarantee.
Yes. Fable leads on Belebele, Global-MMLU, and MMLU-Pro. We show those gaps as clearly as the results where Echo matches it.
The 100-question set was used while building Echo, so it is useful but not an independent test. A separate 41-question set was tested later, but Fable answered only 19 of those questions, so the two overall scores cannot be compared directly. The 19 questions both answered remain visible as an exploratory view.
Math and multiple-choice answers are checked against the known answer. Code is run against tests. We keep the checked answers fixed so the displayed totals do not change silently.
The newest score summaries passed our arithmetic checks, but their question-by-question answers are not yet published here. They remain clearly labeled as summaries until those examples are added.
For developers
Read-only access to the same results shown on this page.
curl {base}api/current
curl {base}api/benchmarks
curl '{base}api/benchmarks/MATH-500/rows?offset=0&limit=50'
curl {base}api/rows/math500:0280d332515ce2d26069
curl {base}api/manifest