Exploratory actual result · 2026-08-28

Seven-case exploratory stack matrix

The operational rows used the same seven cases and scoring fields, but different models and stack configurations. This is stack-level evidence, not a same-model harness attribution test.

Seven-case exploratory stack matrixPass rate for three operational rows on the same seven custom cases. Models and stack configurations differ. Claude was nonoperational and OpenCode was skipped. This matrix does not prove general model or harness quality, market leadership, causal Flywheel advantage, production reliability, or independent validation.Seven-case exploratory stack matrixPASS RATE · ZERO-BASED SCALEFlywheel serve14b-cpt-adapter · 4/7 passed57.1% Ollamaqwen2.5:7b · 5/7 passed71.4% OpenAI Codexgpt-5.3-codex-spark · 0/7 passed0%
Text equivalent for the three operational rows.
StackBackendModelPassed / nPass rateMean qualityMean latencyError rateFailure classes
Flywheel servem7-source-mined-serve14b-cpt-adapter4/757.1%0.63930,926.429 ms0%none: 7
Ollamam7-source-mined-ollama:qwen2.5:7bqwen2.5:7b5/771.4%0.6231,361.571 ms0%none: 7
OpenAI Codexcodex-plangpt-5.3-codex-spark0/70%0.28484,466.286 ms0%low_task_focus: 5; none: 2

Unavailable rows

Systems excluded from numeric ranking rather than represented as zero.
SystemStatusReason
Claude CodeNOT OPERATIONALquota_or_rate_limit; requested-model metadata is not treated as valid
OpenCodeSKIPPEDno configured endpoint backend for provider=opencode modes=plan,api,provider,cloud
Source SHA-256
dcac28e8729d2255f3feb647c76ec00d5a975f8b616e947b92e827daf15002d6
Denominator
7 custom cases per operational row
Environment
not recorded in source artifact
Method
Same seven cases and scoring fields; different harness stacks and different requested models.
Limitations
Exploratory single-run matrix with seven custom cases. The compared rows use different models and stack configurations; this is not a same-model harness attribution test. Hardware, runtime versions, cost, and resource use are not recorded in the source artifact. Claude was nonoperational and OpenCode was skipped, so neither is plotted or scored as zero.

What this does not prove: This matrix does not prove general model or harness quality, market leadership, causal Flywheel advantage, production reliability, or independent validation.