## Education as the ability to judge the work

I expect educational practice to change quickly as AI-assisted work becomes ordinary. The speed and shape of that change remain predictions. Schools, disciplines and learners will experience it differently. Teachers already have substantial knowledge about assessment, feedback and how understanding develops; that knowledge belongs in the design of these systems.

Australia's Tertiary Education Quality and Standards Agency published assessment guidance in 2023 and a 2025 resource on putting those principles into practice. Its principles call for multiple forms of assessment that account for context, alongside ethical participation in an AI-rich society. This is an existing reform effort that the tools can learn from. It does not endorse this portfolio. [TEQSA assessment principles](https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/assessment-reform-age-artificial-intelligence), [implementation guidance](https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/enacting-assessment-reform-time-artificial-intelligence).

A field experiment by Hamsa Bastani, Osbert Bastani, Alp Sungu and their coauthors gives one reason to measure learning separately from assisted performance. In mathematics sessions at one Turkish high school, students using a general GPT-4 interface performed better during assisted practice but worse on a subsequent unassisted exam than the control group. A tutor designed with instructional safeguards largely mitigated that harm. The study measured short-term outcomes in a specific setting; its findings do not establish the effects of every tutor or today's models. [Published study](https://doi.org/10.1073/pnas.2422633122), [author-hosted paper](https://hamsabastani.github.io/education_llm.pdf).

For my own learning tools, I want to ask what the person can explain afterward. Can they locate the source of a claim? When a condition changes, can they adapt the method? Could they recognize a plausible answer that violates the problem's assumptions? These questions should influence the lesson and the assessment from the start.

One proposed exercise begins with a small program and a claim about its behavior. The learner examines the input, predicts the result, then runs it. A second case changes the input so the original reasoning no longer applies. The learner has to explain the difference and identify what evidence would settle a disagreement. Assistance can be declared and adjusted to the lesson. A later task can examine what the learner retained, with accessibility needs accounted for in its design.

That exercise is a design proposal. A convincing demonstration would require educators, a suitable comparison, clear learning outcomes and follow-up beyond the assisted session. The present tool portfolio has no established educational-effectiveness result from such a study.

Records can help a learner show revisions and explain a decision. They can also expose private drafts or create pressure to perform a prescribed style of thinking. Students need clear boundaries around collection and retention, access to their own records and a way to challenge an assessment. A software trace cannot reveal everything a person understood. It can supply evidence for a conversation with someone capable of judging the subject.

This is why learning belongs inside the mission. The design should leave the person better able to use the method, question the tool and continue elsewhere. Notifications and engagement measures need to serve that purpose. Time returned to a student's life can matter more than time retained inside an application.

## Contribution needs a route to correction

For an alignment researcher, the useful starting point may be an evaluation record with a known ambiguity. For a maintainer, it may be a bug that can be reproduced from a small input. I want to offer the relevant component at that scale, with the assumptions and costs visible, so the recipient can decide whether it helps their work.

An evaluation should include an ordinary case the check is expected to accept and a deliberately wrong case it must reject. If a grader gives both a pass, the instrument needs repair before its scores can support a conclusion. The review also has to ask which parts the agent could influence: the submitted artifact, the test environment, the expected answer or the record of the result. A protected external grader is an existing baseline that any proposed improvement must take seriously.

Independent review requires more than a second agent's agreement. Reviewers can share sources, model behavior and incentives. A useful review records what the reviewer could inspect, how the check differed from the original attempt, and which dependencies remain shared. Some questions require domain experts or an independently performed experiment. The software should make those needs visible.

The commercial question deserves the same discipline. A tool can be worth paying for if it reduces the time needed to reconstruct an incident, catches a consequential mistake or makes a necessary workflow usable. A pilot should identify the current process and measure the change against it. Installation, operation and review all cost time. A catalog entry or a sent outreach message supplies no evidence of adoption or revenue.

Safety work carries its own tradeoffs. Faster orchestration may help an evaluator run a useful test, while also making harmful activity easier for someone else. A public agent board can help examine collaboration and can spread misleading instructions. More detailed records help review and increase the amount of material that must be protected. Release decisions need to consider who benefits, who bears the risk and which access boundaries fit the intended use.

## The person who can refuse the result

Pick The Lock for Everyone asks me to examine the barriers my own work creates. Someone should be able to contribute before they know my preferred vocabulary. The product page should explain the task it serves. Installation should state its requirements, and an error should give enough context to recover. If the person chooses another provider or tool, their work should remain available in a usable form.

The ethical commitments extend to authorship and labor. Source material needs an honest account of where it came from and what use was permitted. Reviewers and maintainers contribute skilled work. An AI-assisted draft still needs a person who accepts responsibility for publishing it, responds to corrections and represents its status accurately. A system that helps produce more work should also help account for the burden that work places on others.

Privacy protects the space in which people can think, learn and change. Public scrutiny is most useful at the claim or consequential action that needs an account. It should leave room for protected personal lives and for confidential work that a person has chosen to keep private. When a review cannot proceed without that material, the report should explain the limit and the conditions under which a fuller check could occur.

I am asking researchers, teachers and builders to examine specific parts of this work. Choose a claim whose failure would matter, or a workflow that currently takes too much effort to review. The [tool catalog](https://harperz9.github.io/catalog.html) and linked source projects provide starting points. A bounded reproduction, a correction with its source, or a conversation about a real review problem gives us something concrete to improve.

## Sources, authorship and revision status

The [source notes](https://harperz9.github.io/writing/a-witness-should-not-become-a-ruler/source-notes.json) provide pinned references and excerpts for central claims. The [tool source map](https://harperz9.github.io/writing/a-witness-should-not-become-a-ruler/tool-source-map.json) gives exact character ranges and source values for the pinned public product records. The product explanations link to their public pages and implementation sources. They describe the public portfolio at this edition's review point. Planned connections and research applications retain their stated limits; deployment and use require version-specific checks.

The ethical argument follows the author's [Pick The Lock for Everyone](https://harperz9.github.io/pick-the-lock-for-everyone.html), especially its discussion of review debt, private life and the builder's own accountability. EMET's [specification](https://github.com/HarperZ9/emet/blob/main/SPEC.md) supplies the integrity contract. The education section cites TEQSA's guidance and Bastani and colleagues' study; the author's proposed exercises and expectations are separate from those sources' findings.

Codex prepared this expanded wording from the author's request, published writing, inspected project sources and the cited education material. The author requested the expansion and approved this website essay for publication. The manuscript mentioned here is in preparation; this website essay has not undergone journal peer review.
