Berkeley Function Calling Leaderboard

Berkeley Function Calling Leaderboard · v4-ede5081a24bc · A function-calling benchmark published by its Berkeley maintainers, including single-turn, multi-turn, and agentic tasks.

Model Ranking

Source
#1Kimi-K2-InstructMoonshot AI59.06percentMoonshotai-Kimi-K2-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#2Gemini 2.5 FlashGoogle DeepMind56.24percentGemini-2.5-Flash (FC)unknownRecomputedwinner eligibleThird-party benchmark
#3GPT-4.1OpenAI53.96percentGPT-4.1-2025-04-14 (FC)unknownRecomputedwinner eligibleThird-party benchmark
#4GPT-4.1 MiniOpenAI50.45percentGPT-4.1-mini-2025-04-14 (FC)unknownRecomputedwinner eligibleThird-party benchmark
#5Command ACohere46.49percentCommand A (FC)unknownRecomputedwinner eligibleThird-party benchmark
#6Gemini 2.5 Flash-LiteGoogle DeepMind36.87percentGemini-2.5-Flash-Lite (FC)unknownRecomputedwinner eligibleThird-party benchmark
#7GPT-4.1 NanoOpenAI33.05percentGPT-4.1-nano-2025-04-14 (FC)unknownRecomputedwinner eligibleThird-party benchmark
#8Llama-3.3-70B-InstructMeta31.90percentLlama-3.3-70B-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#9phi-4Microsoft28.79percentPhi-4 (Prompt)unknownRecomputedwinner eligibleThird-party benchmark
#10Llama-4-Scout-17B-16E-InstructMeta28.13percentLlama-4-Scout-17B-16E-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#11Nova 2 LiteAmazonSelected model27.10percentAmazon-Nova-2-Lite-v1:0 (FC)unknownRecomputedwinner eligibleThird-party benchmark
#12Llama-3.1-8B-InstructMeta25.83percentLlama-3.1-8B-Instruct (Prompt)unknownRecomputedwinner eligibleThird-party benchmark
#13Ministral-8B-Instruct-2410Mistral AI11.10percentMinistral-8B-Instruct-2410 (FC)unknownRecomputedwinner eligibleThird-party benchmark

13 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 2, 2026. The model you came from is highlighted.

Methodology and Coverage

Ranking protocoloverall accuracy · percent. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.
Acquisition coverage13 published models from 109 source rows; 91 source identities remain quarantined.

Primary Evidence

Publisher Artifacts

Questions

Berkeley Function Calling Leaderboard FAQs

What does Berkeley Function Calling Leaderboard measure?+

A function-calling benchmark published by its Berkeley maintainers, including single-turn, multi-turn, and agentic tasks. Model Markets classifies it as a tool use benchmark and preserves the publisher's v4-ede5081a24bc release as a distinct comparison cohort.

How are models ranked on Berkeley Function Calling Leaderboard?+

Models are ordered by overall accuracy in percent, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.

Which model currently leads Berkeley Function Calling Leaderboard?+

Kimi-K2-Instruct leads the current verified table with 59.06 percent on v4-ede5081a24bc. This is a benchmark-specific result, not a universal model-quality claim.

How many models have a published Berkeley Function Calling Leaderboard score?+

13 catalog models are published from 109 source rows. 91 source identities remain quarantined rather than guessed.

Can Berkeley Function Calling Leaderboard scores be compared with other benchmarks?+

Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings

Why might a model be missing from Berkeley Function Calling Leaderboard?+

A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.

How current is the Berkeley Function Calling Leaderboard leaderboard?+

The current Model Markets snapshot was observed Sep 2, 2026 from Berkeley Function Calling Leaderboard artifacts. The exact publisher source and each retained result artifact are linked on this page.

Send Feedback