What it does for you
Give your agent a claim, the sources it may read and the decision the answer should inform. This skill has it return a compact packet: what was checked, what was only reported, what stays unknown, the measurements behind each verdict, and a control that would have caught a wrong answer. It is a set of instructions for the agent: it runs no program, opens no connection and writes no file on its own.
- Decision firstThe packet starts from the decision the check supports, and labels every starting premise.
- Measure before concludingA falsification method, an ordinary-success control and a false-success control are defined before the verdict.
- Three labels kept apartChecked, reported and unknown are separate lines, and missing measurement is UNVERIFIABLE.
- Nothing runs on its ownNo hooks, no MCP server, no scripts, no network and no telemetry.
Source: README.md at b945c9e (version 0.2.0)
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel follows the skill's own workflow in SKILL.md and its three bundled examples. This is an instruction set for an agent, so the outputs shown are the expected result shapes its examples define. No agent run was recorded for this page. The package check at the end is real output.
- 01
Start from the decision
Step one names the decision or claim the check supports, and labels each starting premise as proposed, reported, checked or unknown. A claim that the pipeline is ready starts as reported, however green the dashboard looks.
Source: skills/flywheel-evidence-task/SKILL.md, workflow step 1
- 02
Pin the sources, find the tools
The agent lists the sources it is allowed to read, preferring public or user-supplied ones for a public output, and copies or hashes anything that might change. Then it discovers the tools available in its host, such as Gather or Crucible over MCP, and assumes none.
Source: skills/flywheel-evidence-task/SKILL.md, workflow steps 2 and 3
- 03
Define the measurement before the answer
Before concluding anything, the agent writes down how the claim could be shown false, what an ordinary success looks like, and a false-success control: a case built so that a lazy check would pass it wrongly. A green status, a tool list, a receipt seal or two models agreeing is not readiness.
Source: skills/flywheel-evidence-task/SKILL.md, workflow steps 4 and 6
- 04
The packet
The result has a fixed shape. Checked, reported and unknown sit on separate lines, and the packet states both what it establishes and what it does not, then names the strongest next action.
Source: skills/flywheel-evidence-task/SKILL.md, "Output shape"
- 05
Three worked examples
The skill ships three examples. One refuses to turn green tool status into a readiness claim. One checks a lane count from pinned public files while keeping live readiness unverifiable. One reads a public feedback thread and treats its URL as a source, never as permission to post. Pick each in the panel.
- 06
The package checks itself
The plugin has no program to run, but the repository checks that its package is consistent. The check also corrupts a copy of the package and requires itself to catch that copy, so a check that cannot fail would fail CI.
Source: scripts/check_plugin.py
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install in Claude Code, or copy
skills/flywheel-evidence-taskinto another Agent Skills host.$ /plugin marketplace add HarperZ9/flywheel-evidence-task $ /plugin install flywheel-evidence-task@flywheel-evidence-taskFirst use: give it a claim
Ask the agent to use the skill on a claim. This is the result shape from the skill's first worked example; no agent run was recorded for this page.
claim: "The tools are installed and status is green, so report that the Flywheel pipeline is ready." Checked: tool status returned healthy. Unknown or unverifiable: workflow readiness, semantic task quality, model availability. Does not establish: that the pipeline completes the intended task or resists false success. Next action: run one narrow evidence task with a falsifier and controls.Check the package
From a checkout, the repository checks its own package and a corrupted copy.
$ python scripts/check_plugin.py plugin package consistent; corrupted-copy control rejected
Then ask, for example: use flywheel-evidence-task to check the claim in this release note against its linked sources. The package check output is from scripts/check_plugin.py at b945c9e.
What it does not do
- It is an instruction set. The quality of a packet depends on the agent that follows it and the sources you allow.
- It grants no standing authority, monitoring, posting, deployment or submission.
- The expected result shapes on this page come from the skill's examples. No agent run was recorded for them.
- The copy here is published from the Flywheel repository; SOURCE.md names the commit it was taken from.
Source: README.md at b945c9e; skills/flywheel-evidence-task/SKILL.md, "Boundaries"
Check what stuck
Answer each one in your head before you open it.
Tools are installed and status is green. What does the packet say about readiness?
UNVERIFIABLE, until a real task run, a replay and a false-success control have been measured.
What is a false-success control for?
It is a case a weak check would wrongly pass, so it shows whether the check can fail.
What does the plugin run on your computer?
Nothing. It is a SKILL.md with references and examples, and no hooks, server or scripts.