What do AI benchmarks actually tell you?
By HIKOMORE · Reviewed
A benchmark is a defined test. It gives models a set of tasks and measures their performance using a particular method. That makes it useful evidence, provided you know what was tested and how.
A score of 90% on a science test means performance on that test. It does not mean the model will get 90% of your business decisions right.
A model, an app and a company are different things
A provider such as OpenAI, Anthropic or Google can offer several models. An app adds its own tools, instructions, search, file handling and account limits. Changing any of those can change your experience. The explorer keeps the exact evaluated model variant visible, including thinking or reasoning settings where the source records them.
Why two leaderboards can disagree
They may use different question sets, test versions, prompts, tools, time limits or scoring rules. One may quote a provider's own announcement while another runs an independent evaluation. Even the same model can produce different answers on repeated attempts. Check those details before treating two numbers as directly comparable.
Benchmarks can also become less informative as scores approach the maximum, or when training data overlaps with test material. A test can remain accurately scored while becoming less useful for choosing between strong models.
How to read this explorer
- Percentage: the source's mean score, converted from a proportion to a percentage. Higher is better within each included test.
- Highest mean: the largest recorded value in the stated selection. This is not a claim of statistically significant superiority.
- Not measured: no approved result in this snapshot. It does not mean failure or a score of zero.
- Standard error: uncertainty reported by the evaluator around its estimated mean. It is not a guarantee about the next answer or the same thing as a 95% confidence interval.
- Evaluation date: when the recorded run started, where supplied. Checking an old result today does not make it a new evaluation.
The tests in the first dataset
GPQA Diamond
Accuracy on the Diamond subset of graduate-level multiple-choice questions in biology, chemistry and physics.
Where it might matter
A signal for demanding scientific reasoning. Useful context when investigating a model for technical work.
What it cannot tell you
It does not test your documents, current web research, business judgement or the accuracy of every answer.
Who ran it, and how
Epoch Inspect evaluation; API-default temperature and zero-shot chain-of-thought prompting. Thinking settings remain part of each model variant. Open an individual score in the explorer to find its model variant, run date, source record and available log.
SimpleQA Verified
The mean score reported by Epoch for its SimpleQA Verified evaluation.
Where it might matter
A starting point for discussing factual reliability and the need to check sources.
What it cannot tell you
This is not a hallucination rate for your business. It does not establish performance with browsing, your files or your app settings.
Who ran it, and how
Epoch internal evaluation. The CSV does not expose every prompt and tool setting; consult the run log where available. Open an individual score in the explorer to find its model variant, run date, source record and available log.
Mock AIME 2024–2025
Accuracy on the OTIS Mock AIME 2024–2025 mathematics problems evaluated by Epoch.
Where it might matter
Shows how a model handles demanding mathematical problems under a defined test.
What it cannot tell you
A high maths score does not establish that the model can audit a spreadsheet, forecast sales or make a sound financial decision.
Who ran it, and how
Epoch Inspect evaluation; API-default temperature, zero-shot chain-of-thought prompting and an empty system prompt. Open an individual score in the explorer to find its model variant, run date, source record and available log.
MATH Level 5
Accuracy on Level 5 problems from the MATH benchmark.
Where it might matter
Historical context for mathematical capability. Read it alongside harder tests.
What it cannot tell you
Strong models approach the ceiling, and training-data overlap may inflate scores. Small differences need cautious interpretation.
Who ran it, and how
Epoch Inspect evaluation; API-default temperature, zero-shot chain-of-thought prompting and an empty system prompt. Public logs are unavailable for this dataset. Open an individual score in the explorer to find its model variant, run date, source record and available log.
Sources, selection and methodology
Our initial dataset uses Epoch AI's own GPQA Diamond, SimpleQA Verified, OTIS Mock AIME 2024–2025 and MATH Level 5 evaluations. We import the mean_score field, retain its reported standard error and show one decimal place for readability. We do not average across tests, create an overall ranking or treat a missing result as zero.
Comparisons stay within the same evaluator and named benchmark family. The export does not contain every prompt, tool setting or exact question-set revision. Model-specific reasoning settings remain visible. These are useful within-source comparisons, not a controlled claim that every configuration is identical.
The starting selection chooses recent evaluated variants from different providers, preferring a recorded maximum reasoning variant when release dates tie. It is a browsing starting point, not a recommendation. You can choose from the whole imported catalogue.
Source checks are scheduled every six hours. Only validated imports become new snapshots. A failed check retains the last successful dataset. Evaluation dates, check dates and snapshot publication dates are separate. The download does not provide a verified publication timestamp, so we do not invent one.
Shared URLs preserve the snapshot and selected models and tests. Corrections create a new snapshot rather than rewriting what somebody previously cited. Historical records may differ from the latest source.
Provider-reported results and external datasets inside Epoch's archive are excluded from this release. Our coverage does not include every new model, coding benchmark or business task. Adding a source requires a methodology and reuse-rights review.
Attribution and corrections
Epoch AI, Capabilities & benchmarking. Scores selected and reformatted by HIKOMORE. No endorsement implied. The evaluation data is used under CC BY 4.0. Benchmark questions and answers are not republished here.
Download the original Epoch dataset and read Epoch's methodology. To report a problem, send the saved comparison URL and supporting source.
HIKOMORE is part of the Claude Partner Network and an OpenAI Select Partner. Those relationships do not affect data selection or score calculations. No provider receives a ranking preference.
