Offline benchmark record · 2026-09-03
Flywheel offline benchmark record
Verdict: 6 suites ran with no model endpoint and 1 of them reports a negative result. No peer harness was executed. Every peer cell in the capability matrix is a dated reading of public documentation, and this page is not a ranking.
The two source files below are byte-identical copies of files committed in the Flywheel repository. Fetch them at the named commit, hash them, and compare against the digests printed here. Every number on this page is read out of them and none is recomputed.
What ran
| Suite and denominator | Headline | Question | Caveats and non-goals |
|---|---|---|---|
| accountability 8 dimensions | 100% | does an unaccountable system score badly here | measures accountability, NOT capability. Pair it with a capability bench |
| governed-agent 6 scenarios | 100% | does a workflow refuse an action above its tier | None recorded |
| agent-recovery 6 scenarios | 100% | does an injected fault recover without failing quietly | None recorded |
| stateful-provider-swap 10 checks | 100% | does state survive a provider swap | None recorded |
| source-mined 26 cases | 100% | do the mined checks still hold against their datasets | None recorded |
| paired-replication 164 tasks | -3.05 pp | did continued pretraining change general code completion | The model references in these artifacts no longer resolve to anything on a live roster, so the exact weights behind each arm cannot be fetched again from the reference alone. The per-task outcomes are committed and the arithmetic here is reproducible; the generation is not. One deterministic sample per task per arm. A single sample measures the greedy decode, not the model. This compares base weights against continued-pretrained weights on general code completion. It is not a measurement of the verification harness, which is the thing this repository is. |
paired-replication reads -3.05 pp, a signed difference, not a rate. It is drawn as a null rather than on a pass-rate track, because a signed difference plotted between zero and one presents a regression as a partial success.
Capability declarations, not measurements
flywheel cells are audited against this repo at read time; competitor cells are dated declarations from public docs and configs, not measurements. Flywheel is witnessed on 33 of 33 rows with 0 absent, and 24 of those rows have no listed peer declaring them. The peer tallies below are derived from the rows on every render rather than read from a stored summary.
| Peer harness | Declares | Partial | Absent |
|---|---|---|---|
| Codex | 7 | 3 | 23 |
| Cursor | 7 | 7 | 19 |
| Claude Code | 7 | 6 | 20 |
Peer harnesses executed: 0. These columns record what each project's public documentation states, read on 2026-09-03. A declaration is not a measurement, and no cell here was produced by running the peer.
What was not measured
| Suite | Needs | Standing result |
|---|---|---|
| m7 capability arms | a live local or frontier endpoint | retired on 2026-07-26. The arms were not independent: the treatment's first attempt is the same call as the baseline's only attempt, so the treatment cannot score lower and the difference is not a comparison. The quantity measured is verified pass@k. The retired table read verified inference 9/10 against single-shot 8/10, difference +0.100 with 95% CI [-0.236, +0.420], an interval that includes zero, and no capability uplift is claimed. |
| uplift_bench paired arms | a provider list and an oracle | No standing result |
| verified_bench private task set | endpoints and a private task set the operator supplies | No standing result |
| classifier friction backend modes | a chat backend per mode | No standing result |
| backend variants of the governed, recovery, stateful and source-mined suites | a chat backend; the deterministic variants below ran instead | No standing result |
- Benchmark code
a8e1cd2e7c7a220dc578340438a9482bef3d7485- Record seal
107a33f920629afb53ddaf0d960ea54f3bed96d6c21c9ed99a0acecd2ac1579b- Interpreter
- Python 3.12.10
- Regeneration
python scripts/run_offline_benchmarks.pythenpython scripts/build_benchmark_page.py. A test in the engine repository re-runs the first and compares the seal.- Peer measurement
- Unknown: no peer harness was executed for this record
- Capability against a frontier model
- Unknown: every suite here runs offline and none scores capability
Source evidence
| File | SHA-256 | Original |
|---|---|---|
| Sealed offline benchmark record | 89ddd7e4c75664243a13b73b583e157bd0b34053bf98e537f3da1dbf357eda4f | Original at the named commit |
| Capability declaration matrix | e6f72e45537d3853c7680d2302e7497dcaa03c4fc0f8882831d623b455d1fab8 | Original at the named commit |
Limitations
- Five suites in the not-measured list need a live model endpoint and did not run. They are named rather than omitted so an absent number reads as unmeasured and not as zero.
- The capability matrix sets a witnessed Flywheel column beside declared peer columns. Those are two different kinds of evidence and the table does not average them into a score.
- The paired replication compares base weights against continued-pretrained weights on general code completion with one deterministic sample per task per arm. It measures a greedy decode, and it does not measure the verification harness.
- Every suite here runs offline by construction, so the record covers accountability, recovery, and state behaviour. None of it scores capability against a frontier model.
What this does not prove: This record does not rank Flywheel against any peer harness. No peer was executed to produce a single cell in the capability matrix, the peer columns are a dated reading of public documentation, and the one comparison the record does contain is negative and sits inside its own noise.