HarperZ9/flywheel-evidence-taskExplainer, built from commit b945c9eAll repository explainers

flywheel-evidence-task

Turn a claim into an evidence packet that says what was checked and what was not.

What it does for you

Give your agent a claim, the sources it may read and the decision the answer should inform. This skill has it return a compact packet: what was checked, what was only reported, what stays unknown, the measurements behind each verdict, and a control that would have caught a wrong answer. It is a set of instructions for the agent: it runs no program, opens no connection and writes no file on its own.

Source: README.md at b945c9e (version 0.2.0)

Watch

A passing check can still be wrong (2 min 24 s, narrated, captioned). The skill asks for a false-success control before any verdict, because a pass alone does not show the check could fail. Transcript, sources and recall questions.

Video walkthrough: coming with the next release.

How it works, one step at a time

Scroll, or use the step buttons. The panel follows the skill's own workflow in SKILL.md and its three bundled examples. This is an instruction set for an agent, so the outputs shown are the expected result shapes its examples define. No agent run was recorded for this page. The package check at the end is real output.

  1. 01

    Start from the decision

    Step one names the decision or claim the check supports, and labels each starting premise as proposed, reported, checked or unknown. A claim that the pipeline is ready starts as reported, however green the dashboard looks.

    Source: skills/flywheel-evidence-task/SKILL.md, workflow step 1

  2. 02

    Pin the sources, find the tools

    The agent lists the sources it is allowed to read, preferring public or user-supplied ones for a public output, and copies or hashes anything that might change. Then it discovers the tools available in its host, such as Gather or Crucible over MCP, and assumes none.

    Source: skills/flywheel-evidence-task/SKILL.md, workflow steps 2 and 3

  3. 03

    Define the measurement before the answer

    Before concluding anything, the agent writes down how the claim could be shown false, what an ordinary success looks like, and a false-success control: a case built so that a lazy check would pass it wrongly. A green status, a tool list, a receipt seal or two models agreeing is not readiness.

    Source: skills/flywheel-evidence-task/SKILL.md, workflow steps 4 and 6

  4. 04

    The packet

    The result has a fixed shape. Checked, reported and unknown sit on separate lines, and the packet states both what it establishes and what it does not, then names the strongest next action.

    Source: skills/flywheel-evidence-task/SKILL.md, "Output shape"

  5. 05

    Three worked examples

    The skill ships three examples. One refuses to turn green tool status into a readiness claim. One checks a lane count from pinned public files while keeping live readiness unverifiable. One reads a public feedback thread and treats its URL as a source, never as permission to post. Pick each in the panel.

    Source: skills/flywheel-evidence-task/examples

  6. 06

    The package checks itself

    The plugin has no program to run, but the repository checks that its package is consistent. The check also corrupts a copy of the package and requires itself to catch that copy, so a check that cannot fail would fail CI.

    Source: scripts/check_plugin.py

Walkthrough

Install it, run it once, then use the main feature. Each command below is real, and so is its output.

  1. Install

    Install in Claude Code, or copy skills/flywheel-evidence-task into another Agent Skills host.

    $ /plugin marketplace add HarperZ9/flywheel-evidence-task
    $ /plugin install flywheel-evidence-task@flywheel-evidence-task
  2. First use: give it a claim

    Ask the agent to use the skill on a claim. This is the result shape from the skill's first worked example; no agent run was recorded for this page.

    claim: "The tools are installed and status is green, so report that the Flywheel pipeline is ready."
    Checked: tool status returned healthy.
    Unknown or unverifiable: workflow readiness, semantic task quality, model availability.
    Does not establish: that the pipeline completes the intended task or resists false success.
    Next action: run one narrow evidence task with a falsifier and controls.
  3. Check the package

    From a checkout, the repository checks its own package and a corrupted copy.

    $ python scripts/check_plugin.py
    plugin package consistent; corrupted-copy control rejected

Then ask, for example: use flywheel-evidence-task to check the claim in this release note against its linked sources. The package check output is from scripts/check_plugin.py at b945c9e.

What it does not do

Source: README.md at b945c9e; skills/flywheel-evidence-task/SKILL.md, "Boundaries"

Check what stuck

Answer each one in your head before you open it.

Tools are installed and status is green. What does the packet say about readiness?

UNVERIFIABLE, until a real task run, a replay and a false-success control have been measured.

What is a false-success control for?

It is a case a weak check would wrongly pass, so it shows whether the check can fail.

What does the plugin run on your computer?

Nothing. It is a SKILL.md with references and examples, and no hooks, server or scripts.