Latest measured results

Newest evaluations

These are our newest score summaries. Question-by-question examples are still being added, so each card says exactly what is and is not available.

Open every result

Previous detailed results

Every card below lets you inspect the exact questions and answers. A direct comparison appears only when Echo and Fable answered the same questions.

Benchmarks are ordered by performance difference.

Next evaluations

More results are coming.

These cards are placeholders, not claims. Results will appear only after the complete evaluation is ready to publish.

Code

SWE-bench Verified

Real software-repair tasks

No result yet
Reasoning

ARC-AGI

Abstract problem solving

No result yet
Code

BigCodeBench

Practical coding tasks

No result yet

Transparency

How to read these results

These are our own published evaluations, not an independent third-party review. We show the limits so the numbers are not asked to prove more than they can.

What can I inspect?

For the detailed results, you can open every question, the correct answer, and both model answers. The page is read-only and does not rerun either model.

When is a direct comparison shown?

Only when Echo and Fable were tested on the same questions. If one side is missing answers, the page says the scores are not directly comparable.

Can these results prove Echo is always better?

No. They show how the tested systems performed on the questions listed here. Public benchmarks can also appear in model training data, so these results are evidence, not a universal guarantee.

Are losses included?

Yes. Fable leads on Belebele, Global-MMLU, and MMLU-Pro. We show those gaps as clearly as the results where Echo matches it.

Why are there three GPQA test sets?

The 100-question set was used while building Echo, so it is useful but not an independent test. A separate 41-question set was tested later, but Fable answered only 19 of those questions, so the two overall scores cannot be compared directly. The 19 questions both answered remain visible as an exploratory view.

How are answers checked?

Math and multiple-choice answers are checked against the known answer. Code is run against tests. We keep the checked answers fixed so the displayed totals do not change silently.

Why are the latest results separate?

The newest score summaries passed our arithmetic checks, but their question-by-question answers are not yet published here. They remain clearly labeled as summaries until those examples are added.

For developers

Evaluation data API

Read-only access to the same results shown on this page.

curl {base}api/current
curl {base}api/benchmarks
curl '{base}api/benchmarks/MATH-500/rows?offset=0&limit=50'
curl {base}api/rows/math500:0280d332515ce2d26069
curl {base}api/manifest