Exploratory actual result · 2026-08-28
Seven-case exploratory stack matrix
The operational rows used the same seven cases and scoring fields, but different models and stack configurations. This is stack-level evidence, not a same-model harness attribution test.
| Stack | Backend | Model | Passed / n | Pass rate | Mean quality | Mean latency | Error rate | Failure classes |
|---|---|---|---|---|---|---|---|---|
| Flywheel serve | m7-source-mined-serve | 14b-cpt-adapter | 4/7 | 57.1% | 0.639 | 30,926.429 ms | 0% | none: 7 |
| Ollama | m7-source-mined-ollama:qwen2.5:7b | qwen2.5:7b | 5/7 | 71.4% | 0.623 | 1,361.571 ms | 0% | none: 7 |
| OpenAI Codex | codex-plan | gpt-5.3-codex-spark | 0/7 | 0% | 0.284 | 84,466.286 ms | 0% | low_task_focus: 5; none: 2 |
Unavailable rows
| System | Status | Reason |
|---|---|---|
| Claude Code | NOT OPERATIONAL | quota_or_rate_limit; requested-model metadata is not treated as valid |
| OpenCode | SKIPPED | no configured endpoint backend for provider=opencode modes=plan,api,provider,cloud |
- Source SHA-256
dcac28e8729d2255f3feb647c76ec00d5a975f8b616e947b92e827daf15002d6- Denominator
- 7 custom cases per operational row
- Environment
- not recorded in source artifact
- Method
- Same seven cases and scoring fields; different harness stacks and different requested models.
- Limitations
- Exploratory single-run matrix with seven custom cases. The compared rows use different models and stack configurations; this is not a same-model harness attribution test. Hardware, runtime versions, cost, and resource use are not recorded in the source artifact. Claude was nonoperational and OpenCode was skipped, so neither is plotted or scored as zero.
What this does not prove: This matrix does not prove general model or harness quality, market leadership, causal Flywheel advantage, production reliability, or independent validation.