What it does for you
Flywheel's rule is no receipt, no accept, and no learned model on the path that accepts an answer. This repository lets you see that rule work on your own machine in one command: a real coding task runs, a receipt is written, a separate process re-derives it, and a forged copy is caught. It also lays out how the receipt could pair with hardware-attested inference, as a proposal.
- One command
python run_demo.py, with no API key, no GPU and no network. - A separate process re-checks
harness.verify_receiptre-derives the receipt in its own process. - A control that must failThe demo forges the receipt and requires the re-derivation to say DRIFT.
- Check it yourselfEdit any field of the receipt and run the verifier again.
Source: README.md at 8932608 (demo against flywheel-verify 1.5.0)
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. Every line is output from demo/run_demo.py at commit 8932608, run against flywheel-verify 1.5.0 installed from PyPI.
- 01
A real task with a real check
The task asks for
merge_intervals: merge overlapping integer intervals and return them sorted, where intervals that touch, like [1,2] and [2,3], merge into one. The oracle is pytest runningtest_merge_intervals.py.The candidate answer comes from a stub, recorded as
model_ref: stub, so the demo needs no model. What is under test is the receipt path.Source: demo/task/task.json
- 02
Run it and write the receipt
The engine runs the oracle against the candidate. The tests pass, the result is accepted, and the receipt records the task, the exact command, the candidate, the oracle's output hash and the verdict.
Source: demo/run_demo.py, steps 2 and 3
- 03
A fresh process re-derives it
A separate process reads the receipt, re-runs the oracle from the task files beside it, and compares. The recomputed verdict and output hash both match the claimed ones: MATCH, exit 0.
Source: README.md, "Quick start"; flywheel
harness/verify_receipt.py - 04
Forge it, and the re-derivation catches it
The demo flips the stored verdict from PASS to FAIL and runs the verifier again. The output hash still matches, but the recomputed verdict is PASS against a claimed FAIL: DRIFT, exit 1.
A receipt that could not be forged would still say MATCH here. This control is what shows the check can fail.
Source: demo/run_demo.py, step 5
- 05
What it proves, and what it does not
The demo ends by saying what it showed: the result re-derives from its receipt, and a forged result does not. It does not show the task is hard, that the answer is the best one, or anything about a model.
The repository also describes pairing a receipt with hardware-attested inference: attestation shows what ran, and the receipt shows the answer re-derives. That pairing is a proposal and is not shipped.
Source: demo/run_demo.py, step 6; README.md, "The two layers" and "Honest state"
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install the engine from PyPI and clone the demo. Python 3.11 or newer.
$ pip install flywheel-verify pytest $ git clone https://github.com/HarperZ9/flywheel-receipt-demo && cd flywheel-receipt-demo/demoFirst run: the demo
Run a task, write its receipt, re-derive it, then forge it and re-derive again.
$ python run_demo.py RESULT: receipt path verified. Honest MATCH, forged DRIFT.Re-derive the receipt yourself
Verify the receipt in a fresh process from the task files beside it.
$ python -m harness.verify_receipt --receipt out/receipt.json --task-dir task claimed verdict PASS, output_hash a2e1d126b3cd1870 recomputed verdict PASS, output_hash a2e1d126b3cd1870 checks output_hash_matches true, verdict_matches true
Output from demo/run_demo.py at 8932608 with flywheel-verify 1.5.0 from PyPI on Windows. The repository's CI pins flywheel-verify 1.0.0.
What it does not do
- The candidate answer comes from a stub. The demo tests the receipt path; generation is outside it.
- On the shipped benchmark the verified loop shows no accuracy gain over single-shot, and no capability uplift is claimed.
- The pairing with hardware-attested inference is a proposal and is not shipped.
Source: README.md at 8932608, "Honest state"; demo/run_demo.py step 6
Check what stuck
Answer each one in your head before you open it.
Who re-derives the receipt?
A separate process running harness.verify_receipt from the task files.
When the verdict is forged to FAIL, which check fails?
verdict_matches. The output hash still matches, and the recomputed verdict is PASS.
What does hardware attestation add that a receipt cannot?
Evidence of what code ran in sealed hardware. The receipt shows the answer re-derives.