Offline benchmark record · 2026-09-03

Flywheel offline benchmark record

Verdict: 6 suites ran with no model endpoint and 1 of them reports a negative result. No peer harness was executed. Every peer cell in the capability matrix is a dated reading of public documentation, and this page is not a ranking.

The two source files below are byte-identical copies of files committed in the Flywheel repository. Fetch them at the named commit, hash them, and compare against the digests printed here. Every number on this page is read out of them and none is recomputed.

Flywheel offline benchmark record6 suites ran with no model endpoint. 5 report a measured rate and 1 reports a signed difference drawn as a null rather than as a rate. The 33 row capability matrix beneath is a dated reading of public documentation, not a measurement of any peer harness.Flywheel offline benchmark record6 suites, no model endpoint required · 5 measured rates · 1 negative resultEvery number is read out of the sealed record. The negative result is drawn as a null, never on a pass-rate track.accountability8 dimensions100%governed-agent6 scenarios100%agent-recovery6 scenarios100%stateful-provider-swap10 checks100%source-mined26 cases100%paired-replication164 tasks-3.05 pp33 capability rows · the Flywheel column is checked against the repository at read time · peer columns are dated declarations2026-09-03 · no peer harness was executed · not a speed, quality, or market ranking

What ran

One row per suite, leading with the number that answers its question.
Suite and denominatorHeadlineQuestionCaveats and non-goals
accountability
8 dimensions
100%does an unaccountable system score badly heremeasures accountability, NOT capability. Pair it with a capability bench
governed-agent
6 scenarios
100%does a workflow refuse an action above its tierNone recorded
agent-recovery
6 scenarios
100%does an injected fault recover without failing quietlyNone recorded
stateful-provider-swap
10 checks
100%does state survive a provider swapNone recorded
source-mined
26 cases
100%do the mined checks still hold against their datasetsNone recorded
paired-replication
164 tasks
-3.05 ppdid continued pretraining change general code completion

The model references in these artifacts no longer resolve to anything on a live roster, so the exact weights behind each arm cannot be fetched again from the reference alone. The per-task outcomes are committed and the arithmetic here is reproducible; the generation is not.

One deterministic sample per task per arm. A single sample measures the greedy decode, not the model.

This compares base weights against continued-pretrained weights on general code completion. It is not a measurement of the verification harness, which is the thing this repository is.

paired-replication reads -3.05 pp, a signed difference, not a rate. It is drawn as a null rather than on a pass-rate track, because a signed difference plotted between zero and one presents a regression as a partial success.

Capability declarations, not measurements

flywheel cells are audited against this repo at read time; competitor cells are dated declarations from public docs and configs, not measurements. Flywheel is witnessed on 33 of 33 rows with 0 absent, and 24 of those rows have no listed peer declaring them. The peer tallies below are derived from the rows on every render rather than read from a stored summary.

Peer cells across all 33 rows, derived from the declaration file.
Peer harnessDeclaresPartialAbsent
Codex7323
Cursor7719
Claude Code7620

Peer harnesses executed: 0. These columns record what each project's public documentation states, read on 2026-09-03. A declaration is not a measurement, and no cell here was produced by running the peer.

What was not measured

Suites that need a live model endpoint, named so an absent number reads as unmeasured.
SuiteNeedsStanding result
m7 capability armsa live local or frontier endpointretired on 2026-07-26. The arms were not independent: the treatment's first attempt is the same call as the baseline's only attempt, so the treatment cannot score lower and the difference is not a comparison. The quantity measured is verified pass@k. The retired table read verified inference 9/10 against single-shot 8/10, difference +0.100 with 95% CI [-0.236, +0.420], an interval that includes zero, and no capability uplift is claimed.
uplift_bench paired armsa provider list and an oracleNo standing result
verified_bench private task setendpoints and a private task set the operator suppliesNo standing result
classifier friction backend modesa chat backend per modeNo standing result
backend variants of the governed, recovery, stateful and source-mined suitesa chat backend; the deterministic variants below ran insteadNo standing result
Benchmark code
a8e1cd2e7c7a220dc578340438a9482bef3d7485
Record seal
107a33f920629afb53ddaf0d960ea54f3bed96d6c21c9ed99a0acecd2ac1579b
Interpreter
Python 3.12.10
Regeneration
python scripts/run_offline_benchmarks.py then python scripts/build_benchmark_page.py. A test in the engine repository re-runs the first and compares the seal.
Peer measurement
Unknown: no peer harness was executed for this record
Capability against a frontier model
Unknown: every suite here runs offline and none scores capability

Source evidence

Public source files and their digests.
FileSHA-256Original
Sealed offline benchmark record89ddd7e4c75664243a13b73b583e157bd0b34053bf98e537f3da1dbf357eda4fOriginal at the named commit
Capability declaration matrixe6f72e45537d3853c7680d2302e7497dcaa03c4fc0f8882831d623b455d1fab8Original at the named commit

Limitations

What this does not prove: This record does not rank Flywheel against any peer harness. No peer was executed to produce a single cell in the capability matrix, the peer columns are a dated reading of public documentation, and the one comparison the record does contain is negative and sits inside its own noise.