Berkeley Function Calling Leaderboard
Berkeley Function Calling Leaderboard · v4-ede5081a24bc · A function-calling benchmark published by its Berkeley maintainers, including single-turn, multi-turn, and agentic tasks.
Model ranking
Publisher benchmark →Source | |||||||
|---|---|---|---|---|---|---|---|
| #1 | Kimi-K2-InstructMoonshot AI | 59.06 | — | — | Moonshotai-Kimi-K2-Instruct (FC)unknown | Recomputedwinner eligible | Third-party benchmark |
| #2 | Llama-3.3-70B-InstructMeta | 31.90 | — | — | Llama-3.3-70B-Instruct (FC)unknown | Recomputedwinner eligible | Third-party benchmark |
| #3 | phi-4Microsoft | 28.79 | — | — | Phi-4 (Prompt)unknown | Recomputedwinner eligible | Third-party benchmark |
| #4 | Llama-4-Scout-17B-16E-InstructMeta | 28.13 | — | — | Llama-4-Scout-17B-16E-Instruct (FC)unknown | Recomputedwinner eligible | Third-party benchmark |
| #5 | Llama-3.1-8B-InstructMeta | 25.83 | — | — | Llama-3.1-8B-Instruct (Prompt)unknown | Recomputedwinner eligible | Third-party benchmark |
| #6 | Ministral-8B-Instruct-2410Mistral AI | 11.10 | — | — | Ministral-8B-Instruct-2410 (FC)unknown | Recomputedwinner eligible | Third-party benchmark |
6 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Aug 29, 2026.
Ranking protocol
overall accuracy · percent. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.
Acquisition coverage
6 published models from 109 source rows; 103 source identities remain quarantined.
Primary evidence