What it does for you
crucible turns a thesis into claims, each paired with the observation that would refute it. It measures each claim, gives each one a verdict from a pure function of the measurement, and writes a sealed record that anyone can recheck. A claim with nothing to measure is reported as unverifiable, never as a pass.
- Claims with their refutationEach claim states what measured result would prove it wrong. A claim that states none cannot pass.
- A verdict with no model in it
verdict_forcompares a deviation with a tolerance. The same record always gives the same verdict. - A record that resists editsTheses, claims, measurements and verdicts are sealed by SHA-256. A flipped verdict fails the recheck.
- A CI gate
crucible ciexits non-zero when a claim loses standing against a sealed baseline.
Source: README.md at 14fca4f (release 1.5.0)
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel follows the bundled example thesis on binary search, from examples/, through crucible run. Every id, verdict and margin is output from crucible at commit 14fca4f.
- 01
A thesis is a list of claims
The example thesis makes three claims about binary search on 1,024 sorted elements: at most 11 comparisons, at most 3 comparisons, and "more elegant than linear search".
Each claim carries a falsification condition: the measured result that would refute it. The third claim has an empty one.
- 02
Steelman each claim
Before anything is measured, each claim gets the strongest test against it. The default steelman restates the claim's own falsification condition as the challenge. The third claim gets none: with no condition, it cannot be refuted, and the record says so.
Source: src/crucible/steelman.py; README.md, "How a thesis run works"
- 03
Measure: deviation against tolerance
A measurement gives each claim a deviation and a tolerance. The worst case for 1,024 elements is floor(log2 1024) + 1 = 11 probes, so the 11 claim deviates by 0 and the 3 claim deviates by 8. Both use a tolerance of 0.5.
The elegance claim gets no measurement.
- 04
The verdict is a pure function
verdict_forruns a fixed ladder. No falsification condition: UNVERIFIABLE. No measurement, or one bound to a different claim: UNVERIFIABLE. A tolerance that differs from the one sealed with the claim: UNVERIFIABLE, so a verdict cannot be rescued by widening it later.Otherwise the margin is (tolerance minus deviation) divided by tolerance. A margin of 0 or more is MATCH; below 0 is DRIFT. Pick each claim in the panel.
Source: src/crucible/verdict.py,
verdict_for - 05
Seal it into the registry
The assessment goes into a content-addressed registry with a seal over the verdicts and another over the measurements. Each claim carries its own SHA-256, and the registry rejects a second thesis with the same id and a different seal.
The assessment seal includes the run's start time, so a new run writes a new seal. The thesis seal and the claim ids stay the same.
Source: src/crucible/registry.py; README.md, "Highlights"
- 06
Flip one verdict, and the recheck catches it
The bundled demo changes the DRIFT verdict to MATCH inside the sealed assessment, then verifies it again. Verification fails: the edited record no longer agrees with what its seal and measurements re-derive.
Source: examples/demo.py
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install from PyPI. Python 3.11 or newer.
$ pip install crucible-benchFirst run: write the example inputs
The package ships a thesis and its measurements. This writes them into a folder you can read.
$ crucible examples --out crucible-examplesTest a thesis
Run the thesis against its measurements. Each claim is steelmanned first, then measured against its tolerance, and the verdicts are sealed into a registry.
$ crucible run crucible-examples/thesis-binary-search.json --measurements crucible-examples/measurements-binary-search.json --registry .crucible-registry ran thesis bd7404c02eb2036e: 3 claim(s) steelman refutations: 3 MATCH 1 DRIFT 1 UNVERIFIABLE 1Catch a flipped verdict
From a checkout, the demo flips one sealed verdict and rechecks it. The recheck fails.
$ python examples/demo.py counts: MATCH 1 DRIFT 1 UNVERIFIABLE 1 assessment seal 07dafef03f5c..., verified True after flipping a DRIFT to a MATCH, verified False <- caught
Output from crucible at 14fca4f; crucible-bench 1.5.0 is the current PyPI release. The assessment seal line is left out because it changes with each run's start time.
What it does not do
- A verdict is only as good as the measurement fed to it. crucible checks that a measurement binds to its claim; it cannot check that the measuring was done well.
- A claim with no falsification condition is never passed. Opinions like "more elegant" stay UNVERIFIABLE by design.
- The LLM-as-judge measure puts a model at the measurement step. The verdict still comes from
verdict_for, but the deviation is the judge's. - A sealed record shows the verdicts were not edited after the run. It does not show the thesis was worth testing.
Source: README.md at 14fca4f, "Highlights" and "How a thesis run works"; src/crucible/verdict.py
Check what stuck
Answer each one in your head before you open it.
Why does the elegance claim read UNVERIFIABLE and not DRIFT?
It states no falsification condition, so verdict_for stops at the first rung. Nothing could refute it, so nothing can confirm it.
What is the margin for the "at most 3" claim, and what verdict does it give?
(0.5 minus 8) divided by 0.5, which is -15. Below zero, so DRIFT.
A measurement arrives with a wider tolerance than the one sealed with the claim. What happens?
UNVERIFIABLE. The sealed tolerance decides, so a verdict cannot be rescued by widening the bound after the seal.
Why does the assessment seal change between runs when the verdicts do not?
The assessment record includes the run's start time. The thesis seal and the claim hashes stay fixed.