Measured model comparison · 2026-08-28

164-task model pass@1 comparison

Same 164 code-completion tasks, same harness, pass@1, greedy decoding, temperature 0. Flywheel 14B was 3.05 percentage points lower. McNemar p=0.404; the observed difference was not statistically significant at 0.05.

164-task model pass@1 comparisonPass at one under the same harness, greedy decoding, and temperature zero. McNemar p equals 0.404; the difference was not statistically significant at 0.05. This result does not show a statistically significant improvement, general coding superiority, agentic reliability, production readiness, or independent validation.164-task model pass@1 comparisonPASS@1 · ZERO-BASED SCALEBase Qwen 14Bollama:qwen2.5-coder:14b-instruct-q4_K_M · 141/164 passed85.98% Flywheel 14Bollama:flywheel-local-coder-14b · 136/164 passed82.93%
Text equivalent for both model results.
Model roleModel referencePassed / nPass@1
Base Qwen 14Bollama:qwen2.5-coder:14b-instruct-q4_K_M141/16485.98%
Flywheel 14Bollama:flywheel-local-coder-14b136/16482.93%
Source
Public tracked result
Source SHA-256
587da8ea4a04d5e520cc5ed4a83efb6c070fab6c587e0e58b00053cfdee5ed9d
Paired outcomes
9 gains; 14 regressions; 127 both pass; 14 both fail
McNemar test
continuity-corrected χ²=0.696, p=0.404, significant at 0.05: no
Limitations
The base artifact digest is not pinned in this result file. The result covers one 164-task code-completion suite and does not measure agentic tool use. Hardware, runtime version, latency, cost, and resource use are not recorded in this result file.

What this does not prove: This result does not show a statistically significant improvement, general coding superiority, agentic reliability, production readiness, or independent validation.