Methodology
Rows are joined only when both models have results for the same benchmark slug, version, and metric. Directionality comes from the benchmark definition rather than an assumed higher-is-better rule.
Model Selection
Compare two models on common benchmark versions collected from original benchmark publishers.
Interactive Tool
| Result | Value | How to read it |
|---|---|---|
| Common benchmark rows | 2 | Same benchmark, version, and metric for both models. |
| Model A published rows | 14 | Claude Fable 5 |
| Model B published rows | 2 | Claude Haiku 4.5 |
| Benchmark | Version and metric | Claude Fable 5 | Claude Haiku 4.5 | Publisher |
|---|---|---|---|---|
| LMArena Document Arena | document-2026-07-30-011508720696arena_rating · higher is better | 1504.249154 rating | 1423.027241 rating | LMArena |
| LMArena Text Arena | text-2026-09-01-011508720696arena_rating · higher is better | 1494.404018 rating | 1395.246555 rating | LMArena |
Rows come from original benchmark publishers. Open the source and protocol before treating a score difference as a general winner.
Inputs
| Input | How it is used |
|---|---|
| Model A and B | Two catalog versions with published benchmark mappings. |
Rows are joined only when both models have results for the same benchmark slug, version, and metric. Directionality comes from the benchmark definition rather than an assumed higher-is-better rule.
Read the evaluation protocol and source artifact before calling a winner. A result can be publisher-verified without being an independent Model Markets measurement.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Model Selector | Turn workload constraints into a short, inspectable model shortlist instead of a universal best-model claim. |
| Price vs Performance | Screen for models that combine useful published performance with acceptable token economics. |
| Reasoning Models | Shortlist reasoning models while keeping output-token economics and evidence coverage visible. |
Questions
Inspect like-for-like published results without blending incompatible metrics or versions. It returns common benchmark rows with version, metric, scores, direction, publisher, and source recency.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
No common rows means insufficient comparable evidence, not a tie. Scores across different benchmarks are not directly interchangeable. Sampling uncertainty and configuration differences can change apparent ordering. Open the linked canonical records and primary sources before making a production or purchasing decision.