What it does for you
A transcript says what a model believes happened; the terminal state says what happened. These environments score by the second. The first one teaches a model to sort agent-run records into seven terminal verdicts and to decide which runs belong in a quality denominator, so a provider outage or a blocked launch stops showing up as a fake regression. Every row is generated from a reference scorer, so every score can be re-derived.
- Exhaustive, not sampledFive fields with every reachable combination: 324 records.
- Ground truth from a scorerThe dataset is generated by a deterministic reference scorer written first.
- Exclusions are scoredOne reward pays for the verdict and another for getting the denominator right.
- Claims pinned as testsExhaustiveness, rewards, and the integrity rule are all tests.
Source: README.md at 510e84c (mlflow-terminal-state 0.1.1)
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel runs the reference scorer and the two reward functions from mlflow_terminal_state.py at commit 510e84c, with stand-in modules in place of the verifiers and datasets packages, which the scorer itself does not need.
- 01
A run record has five fields
Each record describes one agent run on five axes: what the run did, how the provider answered, what the independent check said, whether the record's receipt is intact, and whether the artifact hashes recompute. Every combination is enumerated: 4 x 3 x 3 x 3 x 3 = 324 records, with no sampling.
Source: environments/mlflow_terminal_state/mlflow_terminal_state.py,
FIELDSandbuild_dataset - 02
Eight rules, in a fixed order
The reference scorer settles the states that must not count as failures first. A blocked or never-launched run is not a run. A timeout reached no answer. A structured provider refusal is not a task failure. Output that never parsed is not a claim. Then integrity: a receipt or artifact mismatch refutes the run even if the oracle passed. Only then does the oracle decide, and an absent oracle is an honest null.
Source: environments/mlflow_terminal_state/mlflow_terminal_state.py,
score - 03
Score some records
Pick a record in the panel. An oracle pass with a broken receipt is refuted, because integrity outranks the oracle. An oracle pass with nothing broken is verified. A run with no oracle is unverifiable and left out of the denominator, and so is a timeout, even one whose oracle failed.
Source: environments/mlflow_terminal_state/mlflow_terminal_state.py,
score - 04
Only 23 of 324 belong in the denominator
Across all 324 records, half are runs that never launched and a quarter timed out. Only verified and refuted runs count toward quality: 23 rows. Counting everything that did not pass as a failure would have put 320 rows in a failure column.
Source: environments/mlflow_terminal_state/mlflow_terminal_state.py
- 05
Rewards pay for the exclusion too
A model answers with a JSON verdict and a denominator flag. The verdict is worth 0.7 and the denominator flag 0.3. For a record whose answer is refuted and counted, pick each model answer in the panel. Prose with no JSON scores 0.
Source: environments/mlflow_terminal_state/mlflow_terminal_state.py,
verdict_reward,denominator_rewardandload_environment - 06
A second environment: write and repair components
verifiers-component-forge asks a model to write or repair parts of an evaluation environment against a written contract. Five families carry 337 cases and 5,058 frozen expectations, scored by running the model's code in a child process against hidden probes. Returning a broken module unchanged scores at most the frozen regression share.
This step is described from the README; its runs need the verifiers toolchain and a model.
Source: README.md, "verifiers-component-forge"; environments/verifiers-component-forge/REPRODUCE.md
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Get it
Clone the first environment and install its tools with uv.
$ git clone https://github.com/HarperZ9/terminal-state-fixtures && cd terminal-state-fixtures/environments/mlflow_terminal_state $ uv venv && uv pip install verifiers pytestScore a record
The reference scorer reads a run record's five fields. An oracle pass with a broken receipt is refuted.
record: returned / ok / oracle pass / receipt mismatch / artifact match verdict refuted, in_denominator trueA run with no oracle
A run with no independent check is unverifiable and stays out of the quality denominator.
record: returned / ok / oracle absent / receipt verified / artifact match verdict unverifiable, in_denominator falseRun the pinned claims
The tests pin exhaustiveness, the rewards and the integrity rule. They were not run for this page.
$ uv run pytest tests/ -q
The scorer and reward values above were computed from mlflow_terminal_state.py at 510e84c with stand-in modules for verifiers and datasets; the tests were not run for this page.
What it does not do
- No scored model run is recorded for the first environment, so nothing here is evidence about how a model performs on it.
- Version 0.1.1 is staged privately on the Environments Hub; it is not a public Hub release.
- The environment teaches a fixed classification contract. Real run records may carry states the five fields do not cover.
Source: README.md at 510e84c, "Environments" and the fixture table
Check what stuck
Answer each one in your head before you open it.
An oracle passed but the receipt does not match. What is the verdict?
Refuted, and it counts: integrity outranks the oracle.
How many of the 324 records belong in the quality denominator?
23: the 4 verified and the 19 refuted.
A model gets the verdict right but the denominator wrong. What does it score?
0.7.