Berkeley Function Calling Leaderboard

Berkeley Function Calling Leaderboard · v4-ede5081a24bc · A function-calling benchmark published by its Berkeley maintainers, including single-turn, multi-turn, and agentic tasks.

Source
#1Kimi-K2-InstructMoonshot AI59.06percentMoonshotai-Kimi-K2-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#2Llama-3.3-70B-InstructMeta31.90percentLlama-3.3-70B-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#3phi-4Microsoft28.79percentPhi-4 (Prompt)unknownRecomputedwinner eligibleThird-party benchmark
#4Llama-4-Scout-17B-16E-InstructMeta28.13percentLlama-4-Scout-17B-16E-Instruct (FC)unknownRecomputedwinner eligibleThird-party benchmark
#5Llama-3.1-8B-InstructMeta25.83percentLlama-3.1-8B-Instruct (Prompt)unknownRecomputedwinner eligibleThird-party benchmark
#6Ministral-8B-Instruct-2410Mistral AI11.10percentMinistral-8B-Instruct-2410 (FC)unknownRecomputedwinner eligibleThird-party benchmark

6 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Aug 29, 2026.

Ranking protocol

overall accuracy · percent. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.

Acquisition coverage

6 published models from 109 source rows; 103 source identities remain quarantined.

Primary evidence

Publisher artifacts

Send feedback