Measured model comparison · 2026-08-28
164-task model pass@1 comparison
Same 164 code-completion tasks, same harness, pass@1, greedy decoding, temperature 0. Flywheel 14B was 3.05 percentage points lower. McNemar p=0.404; the observed difference was not statistically significant at 0.05.
| Model role | Model reference | Passed / n | Pass@1 |
|---|---|---|---|
| Base Qwen 14B | ollama:qwen2.5-coder:14b-instruct-q4_K_M | 141/164 | 85.98% |
| Flywheel 14B | ollama:flywheel-local-coder-14b | 136/164 | 82.93% |
- Source
- Public tracked result
- Source SHA-256
587da8ea4a04d5e520cc5ed4a83efb6c070fab6c587e0e58b00053cfdee5ed9d- Paired outcomes
- 9 gains; 14 regressions; 127 both pass; 14 both fail
- McNemar test
- continuity-corrected χ²=0.696, p=0.404, significant at 0.05: no
- Limitations
- The base artifact digest is not pinned in this result file. The result covers one 164-task code-completion suite and does not measure agentic tool use. Hardware, runtime version, latency, cost, and resource use are not recorded in this result file.
What this does not prove: This result does not show a statistically significant improvement, general coding superiority, agentic reliability, production readiness, or independent validation.