HarperZ9/terminal-state-fixturesExplainer, built from commit 510e84cAll repository explainers

terminal-state-fixtures

Score an AI agent by the state its work leaves behind.

What it does for you

A transcript says what a model believes happened; the terminal state says what happened. These environments score by the second. The first one teaches a model to sort agent-run records into seven terminal verdicts and to decide which runs belong in a quality denominator, so a provider outage or a blocked launch stops showing up as a fake regression. Every row is generated from a reference scorer, so every score can be re-derived.

Source: README.md at 510e84c (mlflow-terminal-state 0.1.1)

Watch

Models do what training pays for (2 min, narrated, captioned). These environments pay a model for the state its work leaves and for the runs it excludes, so the reward measures what was meant. Transcript, sources and recall questions.

Video walkthrough: coming with the next release.

How it works, one step at a time

Scroll, or use the step buttons. The panel runs the reference scorer and the two reward functions from mlflow_terminal_state.py at commit 510e84c, with stand-in modules in place of the verifiers and datasets packages, which the scorer itself does not need.

  1. 01

    A run record has five fields

    Each record describes one agent run on five axes: what the run did, how the provider answered, what the independent check said, whether the record's receipt is intact, and whether the artifact hashes recompute. Every combination is enumerated: 4 x 3 x 3 x 3 x 3 = 324 records, with no sampling.

    Source: environments/mlflow_terminal_state/mlflow_terminal_state.py, FIELDS and build_dataset

  2. 02

    Eight rules, in a fixed order

    The reference scorer settles the states that must not count as failures first. A blocked or never-launched run is not a run. A timeout reached no answer. A structured provider refusal is not a task failure. Output that never parsed is not a claim. Then integrity: a receipt or artifact mismatch refutes the run even if the oracle passed. Only then does the oracle decide, and an absent oracle is an honest null.

    Source: environments/mlflow_terminal_state/mlflow_terminal_state.py, score

  3. 03

    Score some records

    Pick a record in the panel. An oracle pass with a broken receipt is refuted, because integrity outranks the oracle. An oracle pass with nothing broken is verified. A run with no oracle is unverifiable and left out of the denominator, and so is a timeout, even one whose oracle failed.

    Source: environments/mlflow_terminal_state/mlflow_terminal_state.py, score

  4. 04

    Only 23 of 324 belong in the denominator

    Across all 324 records, half are runs that never launched and a quarter timed out. Only verified and refuted runs count toward quality: 23 rows. Counting everything that did not pass as a failure would have put 320 rows in a failure column.

    Source: environments/mlflow_terminal_state/mlflow_terminal_state.py

  5. 05

    Rewards pay for the exclusion too

    A model answers with a JSON verdict and a denominator flag. The verdict is worth 0.7 and the denominator flag 0.3. For a record whose answer is refuted and counted, pick each model answer in the panel. Prose with no JSON scores 0.

    Source: environments/mlflow_terminal_state/mlflow_terminal_state.py, verdict_reward, denominator_reward and load_environment

  6. 06

    A second environment: write and repair components

    verifiers-component-forge asks a model to write or repair parts of an evaluation environment against a written contract. Five families carry 337 cases and 5,058 frozen expectations, scored by running the model's code in a child process against hidden probes. Returning a broken module unchanged scores at most the frozen regression share.

    This step is described from the README; its runs need the verifiers toolchain and a model.

    Source: README.md, "verifiers-component-forge"; environments/verifiers-component-forge/REPRODUCE.md

Walkthrough

Install it, run it once, then use the main feature. Each command below is real, and so is its output.

  1. Get it

    Clone the first environment and install its tools with uv.

    $ git clone https://github.com/HarperZ9/terminal-state-fixtures && cd terminal-state-fixtures/environments/mlflow_terminal_state
    $ uv venv && uv pip install verifiers pytest
  2. Score a record

    The reference scorer reads a run record's five fields. An oracle pass with a broken receipt is refuted.

    record: returned / ok / oracle pass / receipt mismatch / artifact match
    verdict refuted, in_denominator true
  3. A run with no oracle

    A run with no independent check is unverifiable and stays out of the quality denominator.

    record: returned / ok / oracle absent / receipt verified / artifact match
    verdict unverifiable, in_denominator false
  4. Run the pinned claims

    The tests pin exhaustiveness, the rewards and the integrity rule. They were not run for this page.

    $ uv run pytest tests/ -q

The scorer and reward values above were computed from mlflow_terminal_state.py at 510e84c with stand-in modules for verifiers and datasets; the tests were not run for this page.

What it does not do

Source: README.md at 510e84c, "Environments" and the fixture table

Check what stuck

Answer each one in your head before you open it.

An oracle passed but the receipt does not match. What is the verdict?

Refuted, and it counts: integrity outranks the oracle.

How many of the 324 records belong in the quality denominator?

23: the 4 verified and the 19 refuted.

A model gets the verdict right but the denominator wrong. What does it score?

0.7.