Midium

Tool calling, measured.

Gemma 4 26B A4B against Ollama. Same model, matched 4-bit, graded by Berkeley's checker.

90.2%

Tool call accuracy on a 26B model at 4-bit. Ollama’s best is 87.4%. Parallel tool calls take a 13.5-point lead.

Midium tool-call accuracy against Ollama, matched on thinking and no-think, including parallel calls, on Gemma 4 26B A4B.
TestMidiumOllama
ThinkingTool call accuracy90.2%↑ 2.8 Ollama thinking↑ 5.2 Ollama no-think87.4%
No-thinkSame test, thinking off89.7%↑ 2.3 Ollama thinking↑ 4.7 Ollama no-think85.0%
Parallel tool callsThe most difficult tool-calling test
No-think83.0%
↑ 5 Ollama thinking↑ 13.5 Ollama no-think
Thinking78.0%
No-think69.5%

Measured July 23, 2026 · Midium vs Ollama · Gemma 4 26B A4B, matched 4-bit · Apple Silicon, 128 GB · BFCL v4 official checker · 1,000 cases · temp 0.

Median tokens
75
Ollama thinking uses 294
Median latency
1.13s
Ollama thinking takes 3.5s
Decode speed
+37%
123.8 vs 90.1 tok/s
Sooner to first token
2.3×
0.122s vs 0.277s

Methodology

Measured July 23, 2026 on Apple Silicon with 128 GB of memory. Both engines served Gemma 4 26B A4B at matched 4-bit, MLX 4-bit and Ollama q4_K_M, on the native tool-calling path. BFCL v4 official checker, 1,000 cases, temp 0. Midium Cloud serves the same 4-bit quantization measured here.

Read the measurement.