Latest measured results

Newest evaluations

These are our newest score summaries. Question-by-question examples are still being added, so each card says exactly what is and is not available.

Open every result

Recorded question-by-question results

Every card below lets you inspect the exact questions and stored answers. A direct comparison appears only when Echo and Fable answered the same questions.

These scores come from the recorded runs shown here. Each row preserves one run per available system. A new run of the same prompt can produce a different answer.

Benchmarks are ordered by performance difference.

Next evaluations

More results are coming.

These cards are placeholders, not claims. Results will appear only after the complete evaluation is ready to publish.

Code

SWE-bench Verified

Real software-repair tasks

No result yet
Reasoning

ARC-AGI

Abstract problem solving

No result yet
Code

BigCodeBench

Practical coding tasks

No result yet

Transparency

How to read these results

These are our own published evaluations, not an independent third-party review. We show the limits so the numbers are not asked to prove more than they can.

What can I inspect?

For the detailed results, you can open every question, the correct answer, and both stored model answers. If an answer was not returned, the row says so explicitly and leaves it out of direct comparisons. The page is read-only and does not rerun either model.

Can a rerun produce a different answer?

Yes. Model outputs can vary between runs. This page preserves the exact measured run used for each displayed score, and every question view shows the number of stored runs.

When is a direct comparison shown?

Only when Echo and Fable were tested on the same questions. If one side is missing answers, the page says the scores are not directly comparable.

Can these results prove Echo is always better?

No. They show how the tested systems performed on the questions listed here. Public benchmarks can also appear in model training data, so these results are evidence, not a universal guarantee.

Are losses included?

Yes. We show recorded losses as clearly as the results where Echo matches or leads. Evaluations still being completed are withheld until their runs are ready.

Why are there three GPQA test sets?

The 100-question set was used while building Echo, so it is useful but not an independent test. A separate 41-question set was tested later, but Fable answered only 19 of those questions, so the two overall scores cannot be compared directly. The 19 questions both answered remain visible as an exploratory view.

How are answers checked?

Math and multiple-choice answers are checked against the known answer. Code is run against tests. We keep the checked answers fixed so the displayed totals do not change silently.

Why are the latest results separate?

The newest score summaries passed our arithmetic checks, but their question-by-question answers are not yet published here. They remain clearly labeled as summaries until those examples are added.

For developers

Evaluation data API

Read-only access to the same results shown on this page.

curl {base}api/current
curl {base}api/current/MATH-500/evidence
curl '{base}api/current/MATH-500/rows?offset=0&limit=50&outcome=all'
curl {base}api/current/MATH-500/rows/math500:000:13bb5d32a5afacb3
curl {base}api/benchmarks
curl '{base}api/benchmarks/MATH-500/rows?offset=0&limit=50'
curl {base}api/rows/math500:0280d332515ce2d26069
curl {base}api/manifest