Tool calling, measured.
Gemma 4 26B A4B against Ollama. Same model, matched 4-bit, graded by Berkeley's checker.
90.2%
Tool call accuracy on a 26B model at 4-bit. Ollama’s best is 87.4%. Parallel tool calls take a 13.5-point lead.
| Test | Midium | Ollama |
|---|---|---|
| ThinkingTool call accuracy | 90.2%↑ 2.8 Ollama thinking↑ 5.2 Ollama no-think | 87.4% |
| No-thinkSame test, thinking off | 89.7%↑ 2.3 Ollama thinking↑ 4.7 Ollama no-think | 85.0% |
| Parallel tool callsThe most difficult tool-calling test | No-think83.0% ↑ 5 Ollama thinking↑ 13.5 Ollama no-think | Thinking78.0% No-think69.5% |
Measured July 23, 2026 · Midium vs Ollama · Gemma 4 26B A4B, matched 4-bit · Apple Silicon, 128 GB · BFCL v4 official checker · 1,000 cases · temp 0.
- Median tokens
- 75
- Ollama thinking uses 294
- Median latency
- 1.13s
- Ollama thinking takes 3.5s
- Decode speed
- +37%
- 123.8 vs 90.1 tok/s
- Sooner to first token
- 2.3×
- 0.122s vs 0.277s
Methodology
Measured July 23, 2026 on Apple Silicon with 128 GB of memory. Both engines served Gemma 4 26B A4B at matched 4-bit, MLX 4-bit and Ollama q4_K_M, on the native tool-calling path. BFCL v4 official checker, 1,000 cases, temp 0. Midium Cloud serves the same 4-bit quantization measured here.