Benchmarks

Benchmarks

Versioned third-party benchmarks and independent workload tests with source classes kept separate.

BFCL

Tool-calling behavior and correctness. Imported results remain third-party benchmark evidence.

HELM

Transparent multi-metric evaluation scenarios with explicit benchmark versions.

lm-evaluation-harness

Reproducible task evaluations with stored configuration and model version.

SWE-bench

Real-world software issue resolution results, separated from inference-provider telemetry.

No benchmark results are published yet.The schema and ingestion boundary are ready. Results will be added only with reliable model-version mapping, reproducible configuration, and a raw source artifact.

Questions

Benchmark questions

Which benchmarks does Model Markets track?+

The source registry is designed for BFCL, HELM, lm-evaluation-harness, SWE-bench, and other reproducible primary benchmark releases. No benchmark result is shown until its version and provenance are stored.

Are benchmark results independently measured?+

Third-party benchmark imports and Model Markets active probes use different source classifications and are never labeled as the same evidence.

How current are benchmark results?+

Each result records its benchmark version, model version, observation and effective timestamps, source URL, collector version, and raw evidence reference.

Why are there no benchmark scores yet?+

This first delivery establishes the contracts and source separation. It intentionally does not manufacture results or run a large original benchmark suite.

How are latency and throughput measured?+

Active probes will run from a disclosed AWS region and preserve workload, provider, region, service tier, request parameters, and returned model version. Read the methodology