Quick answer
Yes, with caveats. An open model of around 30B running locally handles a few well-described tools and a short loop well: pick a tool, fill in the arguments, read the result, answer. Benchmarks show that single calls are no longer a problem even for 8B models. The problem is long, multi-turn tasks where the model plans the next steps itself: that is where smaller models' scores halve. So an agent on a local model needs argument validation, a step limit, a loop breaker and human approval before actions with consequences. That is how we build agents at pimento, on models from the Qwen, Gemma, GLM and Mistral families served on our own NVIDIA Blackwell GPUs.
What a model has to do in an agent loop
"Tool calling" sounds like one skill, but it is several:
- Deciding whether to call a tool at all. A "good morning" should not trigger a customer database query.
- Choosing the right tool out of several, or several dozen.
- Filling in arguments according to the schema: types, required fields, date formats, identifiers.
- Reading the result and deciding what's next: another tool, a question for the user, or an answer.
- Stopping when the task is done or can't be done.
Points 1-3 are a single call. Points 4-5 are the loop, and errors compound there. A model that gets 90% of single steps right can fail more often than it succeeds on a task with ten dependent steps. We cover what an agent is and how it differs from a chatbot in Agentic AI explained.
What the benchmarks show
The most widely cited function-calling leaderboard is the Berkeley Function Calling Leaderboard (BFCL), now at V4. Besides simple calls (single, multiple, parallel), it measures queries from real users ("live"), multi-turn tasks, recognising when no tool fits, and since V4 also agentic scenarios: web search and memory management.
Selected rows from the leaderboard as of 12 April 2026:
| Model (mode) | License | Overall | Simple calls | "Live" queries | Multi-turn |
|---|---|---|---|---|---|
| Claude Opus 4.5 (FC) | closed | 77.47% | 88.58% | 79.79% | 68.38% |
| GLM-4.6 (FC, thinking) | MIT | 72.38% | 87.56% | 80.90% | 68.00% |
| Kimi K2 Instruct (FC) | modified MIT | 59.06% | 81.60% | 78.68% | 50.63% |
| Qwen3-235B-A22B-Instruct-2507 (prompt) | Apache 2.0 | 52.15% | 90.33% | 78.68% | 44.62% |
| Qwen3-32B (FC) | Apache 2.0 | 48.71% | 88.77% | 82.01% | 47.87% |
| Qwen3-8B (FC) | Apache 2.0 | 42.57% | 87.58% | 80.53% | 41.75% |
| Qwen3-30B-A3B-Instruct-2507 (FC) | Apache 2.0 | 41.39% | 85.77% | 77.94% | 30.00% |
| Llama-3.3-70B-Instruct (FC) | Llama 3 Community | 31.90% | 88.02% | 76.61% | 21.50% |
| Gemma-3-27b-it (prompt) | Gemma Terms of Use | 29.47% | 87.17% | 74.54% | 10.75% |
Three things follow from this table.
Single calls are solved. The simple-calls column is flat: 85-90% for almost everyone, from 8B up to the largest closed models. If your agent is "understand the question, call one tool, write the answer", a small local model is enough.
Multi-turn is where models differ. On multi-turn tasks the spread is huge: from about 10% to about 68%. The best open model, GLM-4.6, matches the best closed one here, but it is a large MoE model. Models that fit easily on a single GPU score 30-48%.
The call mode matters. BFCL tests some models twice: through the native function-calling format (FC) and through instructions in the prompt. DeepSeek-V3.2-Exp scores 34.85% on simple calls in FC mode and 85.52% in prompt mode with thinking. Same model. The difference comes from the format and how it is parsed, which is to say from how the model is served.
One caveat: the leaderboard lags behind releases. The April 2026 version does not include, for example, Gemma 4 or newer Qwen generations, even though both now support tools natively. Treat BFCL as a map of model families, not a verdict on a specific version.
The other important benchmark is τ-bench (Sierra), which simulates an agent talking to a user in retail and airline settings, with real business rules. Its authors introduced the pass^k metric: the probability that an agent solves the same task in k attempts in a row. In the original paper, even GPT-4o solved fewer than 50% of tasks, and pass^8 in the retail domain was below 25%. Its successor, τ²-bench, adds scenarios where the user also has to take actions. The takeaway for deployments is simple: average success hides instability, and in production what counts is whether the agent works every time.
Serving: where things usually break
A common cause of "the model can't call tools" lies in serving, not in the model. The model emits a call in its own format (XML tags, JSON inside special tokens, Python-style syntax), and the server has to recognise it and turn it into the tool_calls field of the OpenAI-compatible API. If the parser doesn't match the model, the call lands as plain text and the agent "does nothing".
In vLLM, tool calling is enabled with --enable-auto-tool-choice and --tool-call-parser, and the parser has to match the model family:
| Model family | vLLM parser |
|---|---|
| Qwen2.5, QwQ, Hermes | hermes |
| Qwen3-Coder | qwen3_xml |
| Mistral | mistral |
| Llama 3.1, 3.2 | llama3_json |
| Llama 4 | llama4_pythonic (recommended) |
| GLM-4.5, GLM-4.7 | glm45, glm47 |
| gpt-oss | openai |
| DeepSeek V3, V3.1 | deepseek_v3, deepseek_v31 |
| Kimi K2 | kimi_k2 |
Parser names change between vLLM versions and new models get new parsers, so check the docs for the version you actually run, and the model card. A few more things worth knowing:
tool_choice="required"and named functions are implemented in vLLM through structured outputs. The docs say plainly that this guarantees a syntactically valid call, not a good one.- Strict mode. With
tool_choice="auto", arguments are constrained by the schema only if the tool setsstrict: true. Without it the model generates the call freely and the server only extracts it from the text. Schemas for strict mode:additionalProperties: false, all fields required, optional fields as nullable. - Parallel calls behave differently across families. vLLM doesn't support them for Llama 3, but does for Llama 4.
- llama.cpp (
llama-server) supports tools when started with--jinja. It has native formats for some families and a generic format for the rest, which uses more tokens. The docs warn against aggressive KV cache quantization (such as-ctk q4_0), because it noticeably degrades calls. - Gemma 4 has its own tool declaration and call tokens in the chat template. Google's function-calling guide makes a point that is easy to lose: the model executes nothing itself, your code does, so your code has to validate the function name and arguments.
For a broader look at memory requirements for serving models, see our guide to how much VRAM a local LLM needs.
How to design an agent for a local model
A large closed model forgives a lot: vague tool descriptions, twenty tools in context, a long plan. A local model forgives less, so the agent design has to take some of the work off it. Principles that work well here:
- Few tools at a time. Every tool in context is another chance for a mistake. If the agent has many, split them into groups and pick the group first.
- Descriptions written for the model. The tool name says what it does. The description says when to use it and when not to. An example of the arguments helps more than a long paragraph.
- Unambiguous schemas. Enums instead of free text, one date format, identifiers instead of names.
- Validation in code. Every argument is checked before execution. A validation error goes back to the model as a readable message, not a server exception.
- Deterministic code where possible. If the order of steps is fixed, the model doesn't plan it. The model fills the gaps that need language understanding.
- Step budget and loop breaker. An agent that calls the same tool a third time with the same arguments gets stopped.
- Human approval before consequences. The agent can read on its own. Sending a letter, changing data or entering a deadline waits for "Approve". No decision means no.
- Audit log and test set. Every step goes into a log. Every new version of the model, prompt or tools runs through a test set built from real conversations, several times over, so you see the instability and not just the average.
Some of these principles are also security. An agent with tools is a prompt injection target, and a local model is no more resistant than a cloud one. More in our posts on AI agent and MCP security and the OWASP Top 10 for LLM applications.
When a local model is enough, and when it isn't
| Scenario | Local 8-32B model | Notes |
|---|---|---|
| Assistant with 2-5 read-only tools (search, fetch, calculate) | usually enough | best cost-to-value ratio |
| Filling a form or ticket against a schema | enough | structured outputs and validation |
| RAG agent that decides whether to keep searching | enough with a short loop | iteration limit, confidence threshold |
| Multi-step process with 10+ tools and planning | risky | split into smaller agents or use a bigger model |
| Actions with consequences without human oversight | no, regardless of model | approval and logging |
If data can't leave the company, the alternative to a small model isn't the cloud, it's a bigger open model on your own server. GLM-4.6 and the large Qwen MoE models are close to the top of BFCL, but need correspondingly more GPU memory. How we match the model and hardware to an agent is described on the AI agents page.
The business context matters too. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027 due to costs, unclear value or weak risk controls. An agent on a local model that does one job well and predictably has a better chance of surviving than an ambitious do-everything agent.
Checklist before deploying an agent on a local model
- List the tools and mark which only read and which have consequences.
- Pick 2-3 candidate models from families that score well on BFCL, and check their model cards for the tool format.
- Run them with the right parser and chat template. Check on a few examples that calls land in
tool_calls, not in the message content. - Build a test set of 30-100 real tasks with the expected flow.
- Run each task several times. Measure the share of tasks solved every time, not just the average.
- Add argument validation, a step budget, a loop breaker and approval for actions with consequences.
- Turn on an audit log for every step and review failures weekly.
- Repeat the tests on every change of model, parser or server version.
Sources
- Berkeley Function Calling Leaderboard (BFCL) V4
- BFCL V4, Part 1: Web Search (agentic categories)
- Yao et al.: τ-bench, A Benchmark for Tool-Agent-User Interaction in Real-World Domains (2024)
- Barres et al.: τ²-Bench, Evaluating Conversational Agents in a Dual-Control Environment (2025)
- vLLM: Tool Calling
- llama.cpp: Function calling
- Google AI for Developers: Function calling with Gemma 4
- Gartner: Over 40% of agentic AI projects will be canceled by end of 2027 (2025)
