Work with me
Evaluation review, harness integration, agent safety review and incident investigation.
Zain Dana Harper, sole proprietorKent, Washingtonzaindharper@gmail.com
I check whether an AI evaluation supports the claim made from it, make tasks and scorers run the same way in a second harness, review agents before an audit, and build incident records from public sources. The work is built to serve evaluator organizations, AI developers, companies shipping agents and public-interest funders. I have no paid client in this field yet, and no pilot, retainer or engagement with any AI developer. Each service links to published work you can open and check.
Services
Evaluation design review
Problem. You have an evaluation plan or a published result and need to know whether it supports the claim: controls, denominators, intervals, contamination exposure, and what the result cannot show.
Deliverable. A written review with line-level comments, a list of missing controls, a pre-registration draft if none exists, and a statement of what each headline claim does not prove.
Time and price. One to two weeks. Quoted per engagement; a single protocol can be a small first scope.
Evidence.
- Benchmark evidence status: 0 same-task comparisons passed the gate, and 6 named baselines are marked not measured.
- Articulate PR #9: a pre-registered fairness gate that failed and kept a ruleset from shipping.
- No Receipt, No Accept: the evidence rule behind the method.
Harness integration and re-runnable scoring
Problem. You need a task, scorer or benchmark to run the same way outside its home stack, with controls that show the scorer separates right answers from wrong ones, and an archive another reviewer can verify.
Deliverable. A working integration (for example through Inspect), deterministic controls, a hashed reviewer packet with run instructions, and a note on what the controls do and do not establish.
Time and price. Two to four weeks. Quoted per engagement. Can start with a smaller first scope that delivers the controls only.
Evidence.
- METR count_odds reviewer packet: the upstream task image, three controls through a pinned Inspect bridge, and a hashed archive. The page states it is not METR endorsement.
- Flywheel PR #234, merged 2026-09-13.
- deepeval PR #2822: a merged fix to an outside evaluation library.
- Crucible clean-room demo: a verdict packet you can re-derive yourself.
Agent safety and pre-audit review
Problem. You ship an AI agent and need to know what it can reach, whether its logs and scorers survive misuse, and what evidence an accredited audit (for example AIUC-1) or an insurer will ask for.
Deliverable. A findings report with reproduction steps, severity and a fix for each finding, a map of the actions the agent can reach, and an evidence pack organized for your auditor. The review prepares for an audit. It does not certify; certification comes from accredited firms.
Scope. Testing runs only on systems you own or hold a license to test, under a signed written authorization that names the targets, the window and the stop conditions.
Time and price. Two to three weeks. Quoted per engagement, including the work of running my tools inside your systems under the signed authorization.
Evidence.
- Security advisories in my own tools: flaws I found and published against my released versions, each with the fixing release. All are self-filed and none has a CVE id.
- Accountable Surface: agent actions gated by explicit grants, with durable journals (source).
- The Sandbox Was Never Just a Box: what an agent can reach, and whether the records of its actions stay trustworthy.
Incident review, from public records or commissioned
Problem. A lab, regulator, newsroom or affected party needs an account of what happened in an AI incident, who knew when, and which claims rest on which evidence.
Deliverable. A dossier that keeps each evidence lane apart (legal allegation, company report, host telemetry, independent analysis, remediation), with numbered sources, a timeline, a right-of-reply record and stated limits. In a commissioned review, the published report discloses the access granted, any redactions and the client's review rights.
Time and price. Two to four weeks. Quoted per engagement. I decline any client that wants editorial control over conclusions.
Evidence.
- Who Knew First: 9 incidents, 288 assessed decisions and 127 numbered sources, built to apply one standard to every lab. A re-check by the same model family led to five corrections dated 1 October 2026, several for drafting that leaned toward Anthropic; the page shows them.
- Five evidence lanes, one OpenAI and Hugging Face incident: 12 sources kept in separate lanes.
Independence and conflicts review
Problem. An evaluator organization or a lab has to describe its third-party evaluation (for example under California SB 53 or the EU GPAI Code of Practice) and wants its conflicts, access terms and publication rights mapped before an outsider does it.
Deliverable. A party-by-party ledger of access, compute, money and publication rights, a gap list drawn from the AI Evaluator Forum letter's minimum conditions, and a draft conflict-of-interest policy.
Time and price. One to two weeks. Quoted per engagement.
Evidence.
- Who Pays the Referees: access, compute, money and publication rights for frontier AI evaluators, set out party by party.
- My own independence policy, applied to this practice.
Recurring monitoring brief
Problem. Your team needs a dated, checkable record of what frontier labs and safety institutes announced, and what each announcement can and cannot show, without staff reading every release.
Deliverable. A brief on an agreed cadence, weekly or every two weeks. Each edition separates reported claims from what they cannot show and carries a SHA-256 of its canonical record.
Time and price. Ongoing, 30 days' notice to end. Quoted per month for the agreed cadence.
Evidence.
- Frontier Safety briefing: dated editions, each with the hash of its canonical record.
How prices are set
- Every engagement gets a scoped quote once the work is offered. I do not publish an hourly rate, and I do not discount.
- The quote covers labor and time, direct costs and resources, compute, my own tooling, and the work of running and administering that tooling across your existing infrastructure and systems to evaluate or audit them.
- Each quote states the scope, the deliverables, written acceptance criteria and what falls outside the scope. If the budget is smaller, the scope gets smaller.
- No fee, bonus or renewal depends on what a finding says.
Before you engage
- I work as Zain Dana Harper, a Washington sole proprietor operating under my own name. Engagements are contracted with me directly.
- Every engagement runs under a written engagement letter and my independence policy, which sets out my interests, funding, how I declare conflicts and when I decline work.
- Every payment from this work, including API credits, is listed with its exact source in my public income ledger.
- I build with Anthropic and OpenAI models and publish investigations that involve both companies. The policy discloses this and commits to one standard for every lab.
Work I do not take
- Extracting hidden chain of thought through jailbreaks or prompt injection, decrypting reasoning, or bypassing any access control. This holds for every client and for my own tools.
- Offensive testing of any system the client does not own or hold a license to test, or without written authorization.
- Certification or attestation under ISO/IEC 42001, AIUC-1 or NYC Local Law 144. I can prepare evidence for an accredited certifier or auditor. I cannot issue the certificate.
- Legal advice. Compliance mapping is technical evidence for your counsel.
- Work where payment depends on the finding, or where the client controls the conclusions.
Case studies from published work
Each case comes from my own published work. None is a client engagement. Each one states its limits.
Running an outside evaluation task through a second harness
Problem. Does a task built for one stack score the same way in another, and does the scorer separate right answers from wrong ones at all?
Method. I built the upstream METR count_odds task image, ran it through a pinned Inspect bridge and imported the results into Flywheel. Three deterministic controls checked the scorer: a correct answer, a wrong answer and no submission. The run went into a reviewer archive with a SHA-256 hash and a bundled verifier, so a second reviewer can check its manifests and hashes.
Result. The controls scored as expected: correct 1, wrong 0, no submission 0. The integration merged as Flywheel PR #234 on 2026-09-13. The archive is 178,203 bytes, and its hash is on the packet page.
Limits. One task. The controls show that the scorer and bridge behave on known inputs; they measure no model capability. The no-submission control scores 0 through the upstream scorer's fallback, which the packet notes limits what that control proves. The bundled verifier checks manifests and hashes; it does not rerun METR infrastructure. The page states it is not METR endorsement. No outside reviewer has reported checking it.
A fairness gate that failed and stopped a release
Problem. A 2023 article (Liang et al., Patterns 4, 100779) reported that AI-text detectors flag plain English, common in second-language writing, as machine text. Articulate, my writing-quality tool, has rules that can block a draft. Do those rules treat English learners unfairly?
Method. I wrote the fairness gates down before the run, then ran a confirmatory test on PERSUADE 2.0: 14,797 essays by US students in grades 6 to 12, 1,330 of them by recorded English learners. A release check refuses to publish any ruleset that fails a gate, with no override.
Result. The ruleset failed three of the pre-registered gates (G1, G2 and G4); 88 of 152 gate rows passed. The 12 strict profiles blocked 6.0 percent of learner essays and 14.0 percent of other essays, a gap of -8.0 points (95% interval -9.3 to -6.5). The gap ran against the reference group, and G1 tests both directions, so it failed. G4 failed because 31 to 504 essays per profile changed findings when rewrapped. The default profile blocked none in either group. The new ruleset did not ship. Releases since then, through Articulate 0.7.0 on 2026-10-01, keep the published 0.5.2 detector rules.
Limits. An earlier pass on the corpus the rules were tuned on passed every gate, which is why the PR labels it exploratory. PERSUADE 2.0 covers US school writing only. The result says nothing about the tool's accuracy on other writing.
Finding and disclosing a defect in my own verifier
Problem. Crucible returns MATCH or DRIFT verdicts on sealed measurements. A verifier that can report MATCH without support gives false confidence to everyone downstream.
Method. I found the defect in Crucible's verdict path, fixed it in a patched release, and published a GitHub security advisory that names the affected versions and the upgrade.
Result. Advisory GHSA-49qx-cj4f-wfqv (medium, published 2026-09-26): in crucible-bench 1.2.0 and earlier, the tolerance that decides MATCH was not sealed, so a measurement file could widen it after the claim was registered, and every later check agreed. A readiness command also reported MATCH for checks it never ran. Version 1.3.0 seals the tolerance and returns UNVERIFIABLE on a mismatch.
Limits. The advisory is self-filed and has no CVE id. No outside party found or confirmed it. A disclosed defect shows the review habit; it does not show that the current release is free of others.
A cross-lab incident record from public sources
Problem. After an AI incident, the party holding the logs usually decides when the public hears about it. Readers need to know who knew first, and on what evidence, with one standard for every developer.
Method. I built a record from public sources only: filings, company posts, host telemetry, independent analyses and press. The method's aim is to assess every decision in every incident against the same criteria. Sources are numbered. The page discloses that the record was compiled with Claude, and that Anthropic, which makes Claude, is one of the developers in it.
Result. 9 incidents, 288 assessed decisions and 127 numbered sources, by the page's own counts. A companion dossier on the OpenAI and Hugging Face incident keeps five evidence lanes apart across 12 sources. The page carries a dated 1 October follow-up and five corrections dated the same day.
Limits. Public records undercount. Incidents nobody disclosed are absent, so per-lab counts are floors and cannot rank one lab's conduct against another's. The model-assisted compilation is disclosed but has not been replicated by anyone else. On 1 October 2026 a re-check by the same model family that helped compile the record applied a swap test to the Anthropic items on the page and in its working files. It found nine wordings or selections that read more favorably to Anthropic and no line on the page that was false; what tilted was selection, not facts. Several of the five dated corrections fix that drafting. About 900 Anthropic lines in the ledgers and rubric were not read in full, so whether the rubric scores every lab alike is unknown.
Reporting a benchmark comparison as not measured
Problem. Comparisons between coding agents are easy to assemble from mismatched runs. A buyer needs to know which comparisons rest on the same task under the same conditions.
Method. I screened every benchmark-like record against a published scorecard gate for same-task comparison. Each named baseline without a qualifying run is marked not measured, which the page states is different from zero.
Result. 0 same-task comparisons passed the gate. Six named baselines (Aider, Claude Code, Cline, Codex, OpenCode and OpenHands) are marked not measured, and 112 benchmark-like records were kept out. Two scoped comparisons are published separately on the model pass@1 page.
Limits. This is an honest null. It shows the reporting discipline and claims no performance result for Flywheel or any baseline.
What you will not find yet
No paid client, testimonial or third-party review of any artifact exists yet. The six named baselines on the benchmark page are unmeasured. The closest substitute for a reference is evidence you can check yourself: the METR packet, the Crucible demo and the hashes on each Frontier Safety edition. The full reading list is on the publications page.
Contact
Email zaindharper@gmail.com with the problem, the system or report in scope, and the decision the work should inform. I reply with scoping questions, or with a decline that names the rule it falls under.