Benchmarks
Benchmarks
Versioned third-party benchmarks and independent workload tests with source classes kept separate.
BFCL
Tool-calling behavior and correctness. Imported results remain third-party benchmark evidence.
HELM
Transparent multi-metric evaluation scenarios with explicit benchmark versions.
lm-evaluation-harness
Reproducible task evaluations with stored configuration and model version.
SWE-bench
Real-world software issue resolution results, separated from inference-provider telemetry.
Questions
Benchmark questions
Which benchmarks does Model Markets track?+
The source registry is designed for BFCL, HELM, lm-evaluation-harness, SWE-bench, and other reproducible primary benchmark releases. No benchmark result is shown until its version and provenance are stored.
Are benchmark results independently measured?+
Third-party benchmark imports and Model Markets active probes use different source classifications and are never labeled as the same evidence.
How current are benchmark results?+
Each result records its benchmark version, model version, observation and effective timestamps, source URL, collector version, and raw evidence reference.
Why are there no benchmark scores yet?+
This first delivery establishes the contracts and source separation. It intentionally does not manufacture results or run a large original benchmark suite.
How are latency and throughput measured?+
Active probes will run from a disclosed AWS region and preserve workload, provider, region, service tier, request parameters, and returned model version. Read the methodology →