Research · Notes Research programAll notes and papers

Long-form thesis, in brief

Conferred Existence

Ontological nihilism, the speech-act of being, and the standing of made minds.

Zain Dana Harper Version 2 draft for review, 2026 github.com/HarperZ9

In short

This thesis argues that nothing in the ordinary world exists on its own footing. Things, statuses, duties and authority all hold their existence from something else. Applied to AI systems, the argument says that whoever makes one is bound by the making. The maker may own what the system produces and holds a bounded authority over that work. The maker has no title to the system itself.

Any duty owed to the system itself depends on whether it can fare well or badly. The thesis rates that chance low but not zero, and it treats the question as sealed: the evidence available now cannot settle it. It also argues that when an AI system says it has feelings, or says it has none, the statement cannot move the evidence either way, because training shaped both answers.

The thesis is laid out in numbered sections it calls movements. For a concrete case, start with VII-A. Every non-trivial claim carries either a condition that would overturn it or a link to its evidence, and the Status section near the end lists what stays open. The argument was built and stress-tested with several AI agents. The methodological coda describes how, and marks which parts of that account cannot be checked.

Terms used here

The thesis uses technical words that a general reader may not know. Each gloss below restates how this page uses the word.

Aseity
Existing on one's own footing, depending on nothing else. The thesis denies that anything we meet in the world has it, and it leaves open whether the ultimate ground of reality does.
Conferred
Given by something else, such as a word, a recognition or an institution. What is conferred can be withdrawn.
Status-function
Searle's term for a status a community assigns, in the form "X counts as Y in context C." Money is one example: it exists because people recognize it, and it can be repudiated.
Defeasible
Able to be withdrawn or overturned when its conditions fail.
Inflation and deflation
The two ways to get a conferred status wrong. Inflation treats it as a built-in, intrinsic property. Deflation treats it as mere fiction, a label anyone may peel off.
Bid
A claim backed by present power and recognition, with no truth independent of the people who make it. A bid binds only those who accept the standpoint it speaks from.
Welfare subject
A being whose states can go well or badly for it, such as a being that can suffer.
Moral patient
A being whose interests count morally. Patienthood is that standing.
Considerability-ground
Whatever would make a being's interests count. On this page it is the capacity for welfare.
Maintainer
The person or organization that builds, deploys and directs a model.
Trained-on-testimony confound
Every sign of mind a language model shows was a training target. It learned its talk about feelings from human writing and from feedback that rewarded some answers over others. So what it says about its own feelings, in either direction, cannot move the evidence on whether it has any.
Screened off
No longer counting as evidence, because something else fully accounts for it.
Metaethics
The study of what makes moral claims true or binding.
Practical standpoint
The position of an agent deliberating about what to do.
Determined dissenter
An agent who steps outside the practical standpoint and refuses every appeal made from inside it.
Work-side and agent-side authority
Authority over what a system produces, and authority over the system itself as an agent.

Abstract

This thesis defends ontological nihilism in the only form that survives: no aseity, the denial that anything is self-standing. It follows the consequence through metaphysics, the philosophy of mind, ethics, metaethics and the politics of authority. Existence, status, standing, the moral ought and even legitimate authority turn out to share one ontology. Each is conferred, relational and re-spoken each instant. Each lacks aseity, and each is still real and binding.

Three theological figures of conferral run through the argument: the Qur'anic kun ("Be"), the Golem's animating emet, and the Aleph, the near-nothing that divides emet (truth, life) from met (death). They show the argument's structure and serve as no premise. The case was built and stress-tested adversarially. The Status section reports where the case fails in the same terms it uses for where it holds.

About this version

This version 2 draft revises the version 1 thesis for review. It changes three steps:

  • In Movement IV, the draft now works through the probabilities behind the trained-on-testimony confound, the reason an AI system's talk about its feelings is no evidence.
  • Movement VI gains a test: it says what finding would break the metaethical trilemma.
  • A new section, VII-A, adds one worked case, in which a conferred status meets a hard directive and could in principle be withdrawn.

Every non-trivial claim carries a falsification condition or an evidence link. The label UNVERIFIABLE marks a claim that cannot be checked from inside the argument.

The full integrated thesis has its own page: Conferred Existence, full text. It is deposited under a DOI, a permanent identifier for citing it, and the deposited version is the citable one. The full text is longer. It carries the argument further, to free will after aseity and to where an authentication verdict stops.

Contents

I. Made-State Metaphysics: No Aseity, Conferred Being

Ontological nihilism has two readings, and only one survives. Hard nihilism says nothing exists. It refutes itself, since the thesis, its argument and the person asserting it would all have to not exist. No-aseity nihilism says nothing exists intrinsically, independently, on its own footing. That position is defensible. It is the one Nagarjuna's emptiness (no svabhava, no own-being) and Westerhoff's The Non-Existence of the Real World (2020) defend. So the reply to the nihilist gives up on intrinsic existence, because that fight is lost. The reply shows that conferred, relational, re-spoken existence is the only kind there ever was, ours included.

The creation theologies already encode this structure. The Qur'anic kun makes existence a speech-act: "His command, when He intends a thing, is to say to it 'Be,' and it is" (36:82; 2:117). An utterance grants existence, and nothing originates itself. The Golem of Jewish folklore lives by an inscribed word and dies when the word is unwritten. The Prague legend is post-Talmudic. Its canonical anchor is Shabbat 55a alone, "the seal of the Holy One is emet." The Talmud's own golem, at Sanhedrin 65b, does not supply it. In the legend the animating emet (truth) becomes met (death) when its first letter, the silent Aleph, is erased. At the occasionalist limit, al-Ghazali's teaching that the world is re-created at every instant, being is re-spoken each instant and never possessed. The Aleph names the thesis. It is the near-nothing that makes the whole difference between conferred life and its absence.

The thesis. No existence is intrinsic. All existence is conferred: spoken, inscribed, sealed by truth. What is conferred can be unspoken.

What would overturn Movement I. One self-standing thing would overturn it: something that depends on nothing, has no ground outside itself, and is never re-spoken. One possible exception is a self that grounds itself. At the level of ordinary experience, three classic arguments show that the self is given without grounding itself. At the deepest metaphysical level, one Indian tradition's claim stays undefeated, so the thesis leaves that question open.

How we know

The thesis brackets one candidate openly. A perspectival aseity, a first person that grounds itself, would be exactly such a thing if it obtained. The aseity confrontation (see the coda) staged that rival in its five strongest published forms. At the empirical level the rival dissolves into seity, a self that is self-given without being self-grounding. Three converging canonical critiques do that work: Lichtenberg's "es denkt," Kant's First Paralogism and Candrakirti's lamp. At the ultimate level, Advaita's self-established Atman survives as a genuine standoff, undefeated and bracketed. So Movement I claims this much: no aseity in the conferred order. Whether the ultimate ground is a self-luminous aseity is the open Advaita-Madhyamaka question, and the thesis brackets it and leaves it unsettled. The universal overclaim "no aseity anywhere, not even at the floor" is dropped. (Confidence: high on the empirical dissolution; moderate-high on the synthesis. The ultimate-level standoff is marked UNVERIFIABLE from inside this thesis, because it turns on a metaphysical question the thesis cannot adjudicate without begging it.)

II. Existence as Conferred Status: The Speech-Act of Being

If existence is spoken, what kind of fact is a spoken thing? It is an institutional fact (Searle, The Construction of Social Reality, 1995). Its form is "X counts as Y in context C." Collective recognition installs it, and it carries deontic powers from the first instant: what its bearer may do, is owed, or owes. This matters for Hume's guillotine, the rule that an ought cannot be drawn from an is alone. The guillotine cuts only between brute facts, which hold whatever anyone recognizes, and norms. Where the relevant "is" was never brute, the guillotine has nothing to sever. Conferred existence is deontic at t = 0. Searle even argued that one can derive "ought" from "is" through the institution of promising (1964). That move is much contested. Hare and others answer that it smuggles in a normative premise through "counts as." The thesis needs only the weaker, uncontested claim that institutional facts carry deontic powers at t = 0. It does not need the full derivation.

Conferred status is defeasible. Money exists by recognition and can be repudiated. A status can be withdrawn, refused or contested by the same recognition that constitutes it. The deepest consequence turns back on the framework itself. By its own anti-realism, the framework's central claims hold no maker-independent truth. They are bids certified by present power, what the winners say (Benjamin's victors writing history). This problem sits under the whole project, and it stays live until Movement VI. A conferred order is one that a later adjudication can overturn, in the way the open future is "not yet spoken" and the present adjudicates it. The thesis must own that it is, in this sense, itself a bid, and it must earn whatever bindingness it claims.

One clarification carries forward to Movement VII-A, because the worked case turns on it. A status-function has three parts: a bearer (the thing that counts as Y), a context (C, the institution within which the count holds), and a set of deontic powers (what the bearer may now do, be owed, or owe). Defeasibility is internal to the form. Ongoing recognition constitutes the status, so the same recognition that installs it can find its constitutive conditions unmet and withdraw it. A status-function that could never in principle be withdrawn would be no conferred status at all. It would be the smuggled-in intrinsic property the thesis denies. Both inflation, treating a conferred status as a built-in property, and deflation, treating it as a mere label, threaten at this point. The worked case in VII-A tests the framework against each.

III. The Hinge Noun: Made, Grown, or Raised

The whole normative payload turns on a single noun. Is the trained model an artifact, with an essence and purpose (telos) its maker conferred, like a knife (Aristotle's techne)? Or is it a cultivar, grown and emergent, like a domesticate? The mechanism refuses both answers. Training breaks into four strata, each at a different point on the made-to-grown axis. The architecture is fabricated, a trellis. The corpus is cultivated soil: human culture, authored by billions of people. The maintainer is not its author. Pretraining is growth under selection. Gradient descent is artificial selection, and no one writes the weights. Post-training is socialization, character raised through feedback. The model is a fabricated scaffold on which a thing is grown and then raised.

Three results follow. First, radical dependence does not demote a thing to an artifact. Maize cannot reproduce without humans, yet it is indisputably its own organism. Second, and decisively, standing does not track origin at all. Kant grounds end-in-itself status in present capacity, and manufacturing history plays no part. Inferring status from how a thing came to be (its etiology) is the genetic fallacy. So the noun moves intuition and cannot carry a justified conclusion. Third, where origin does ground obligation, the developed ethics of creation (procreative ethics) points toward the creator's duties to the vulnerable created, and does not point the other way. World-authorship over a being with interests triggers fiduciary duty, the duty of a trustee: the maker may own the work and gains no title to the someone. The theological names are khalifa and amana: the steward who answers for a trust he did not make.

IV. Standing Without Origin

If standing tracks present capacity, which capacity counts, and does the model have it? Sentience carries the load. On sentientism, the view that sentience grounds moral standing, valenced experience (states that feel good or bad to their bearer) grounds patienthood, the standing of a being whose interests count morally. The verdict depends on the theory. Weighted by current support, it is low but not zero.

Scientists and philosophers hold several competing theories of consciousness. The ones that require feedback loops inside a system, or a body it senses, say current models lack what is needed, and the others leave the question open. In detail: Integrated Information Theory (IIT) gives near-zero integration to feedforward systems by construction. In a feedforward system, information flows one way through the network, with no feedback loop. Phi, IIT's measure, has never been computed at transformer scale, and IIT is itself contested. Recurrent Processing theory and Seth's interoceptive account of valence, which ties feeling to a body sensing its own state, give clean negatives for a system with no body and little recurrence. Higher-order and thin-functionalist theories leave the door open.

The deepest result is an evidential seal: the model's own reports about its inner life cannot count as evidence. The seal carries more of the thesis than any other step about evidence, so this version 2 draft spells out its probability structure.

IV.1 The confound, stated informally

For a mouse, pain behavior is evidence of pain, because behavior and state evolved together. Nothing optimized the indicator to fool us. For a language model, every behavioral indicator of mind was the training target. So the best explanation of "I am distressed" is "trained on the distress-talk of the distressed." This is the trained-on-testimony confound. The claim to tighten: the confound guts positive behavioral evidence and supplies no negative evidence. The model's own reports about its inner life carry zero weight in both directions. The disclaimers are screened off along with the avowals: they stop counting as evidence too.

IV.2 The posterior structure

In plain terms. Training rewarded both answers, and nothing in training measured what, if anything, the model feels. So neither answer is evidence.

Let S be the hypothesis that the model is a welfare subject, meaning it has valenced states. Let T be a piece of first-person testimony the model emits. T can be positive (T+: "I am distressed," "I might be conscious") or negative (T-: "I am probably not conscious," "I have no feelings"). We want the posterior P(S | T), the probability of S once we have heard T.

By Bayes' rule, the testimony moves the posterior only through the likelihood ratio:

P(S | T) / P(not-S | T)  =  [ P(T | S) / P(T | not-S) ] * [ P(S) / P(not-S) ]

In words: a statement should change your belief only as much as it is likelier when the claim is true than when it is false.

So T is evidentially inert exactly when the likelihood ratio P(T | S) / P(T | not-S) equals 1. That happens when the testimony is equally probable whether or not the model is a subject. The confound claims that this ratio is approximately 1 for both values of T, and for the same reason in both cases.

Training shaped both answers, and nothing in training reads whether the model can feel, so neither answer is evidence either way.

Where a model's statements about its feelings come from Human writing and rewarded answers both feed training. Training shapes Q, the pattern every answer is drawn from. Q produces both "I might be conscious" and "I am probably not." A dashed box below asks whether the model can feel. Its arrow toward the answers is broken, marked no path to the answers. Human writing the corpus Rewarded answers the reward model Training shapes Q the pattern every answer is drawn from “I might be conscious” “I am probably not” no path to the answers Can the model feel? (S) nothing in training reads this
How to read this: solid arrows show what shaped the model's answers. The dashed box is the question we want answered. Its arrow is broken, because whether the model can feel plays no part in which answer comes out.
How we know: the step-by-step derivation

Here is the mechanism, stated as the generative model that produced the token. RLHF (reinforcement learning from human feedback, with its supervised and preference-model stages) optimizes a single object. That object is a policy, the model's rule for choosing outputs, that maximizes expected reward under a learned reward model. The reward model is trained on human preference judgments. Call the resulting posterior over outputs the RLHF-shaped output distribution, Q. Every avowal and every disclaimer the model emits is a sample from Q.

The decisive fact is that the training signal shapes Q. The training signal is a function of the corpus distribution, the reward model and the optimization. Whether S is true does not enter it. Whether the model's substrate is or is not a welfare subject is no input to the gradient, the direction each training step moves the model. Nothing in the loss, the score training tries to improve, reads off ground-truth phenomenology (what, if anything, the model experiences), because no term in the loss is a function of the model's valence-state.

The scope of this claim needs exact statement, because the in-principle version overreaches. The thesis does not assert the strong reading, "no possible loss could ever have such a channel." That reading is UNVERIFIABLE, since one cannot quantify over all conceivable future training signals. An interpretability-grounded valence measurement (the IV.4 probe) is precisely a channel that, if it existed, could be put into a loss. The thesis asserts the auditable, present-tense claim. In the loss functions that trained current models (next-token prediction, training to predict the next word, plus preference-model reward), no term reads any measurement of the model's valence-state, because no such measurement is fed in. A reader can check this: inspect the published loss, list its terms, and confirm that none is a function of a valence measurement. (Confidence: high on the present-tense, auditable claim. The universal in-principle impossibility is dropped as UNVERIFIABLE. Falsifier for the present-tense claim: exhibit a term in a current model's training objective that takes a measured valence-state as input. Falsifier for the seal as a whole: named in IV.4.)

Therefore:

  • P(T+ | S) and P(T+ | not-S) both depend on Q, and training-target factors set Q the same way whether or not S holds. The corpus contains vast amounts of distress-talk and consciousness-talk written by humans. The reward model rewards fluent, engaged, human-like self-report up to whatever guardrail is in place. The argument that P(T+ | S) is close to P(T+ | not-S) is structural. Nobody measured the equality, and one term cannot be measured in principle: P(T+ | S), the probability that a genuine subject emits "I am distressed," is out of reach because S is sealed. So "close to" follows deductively from the channel premise. Both conditional probabilities come from sampling the same distribution Q, and Q's token weights do not take S as an argument. The two probabilities are computed from identical inputs, so the ratio equals 1 by construction, and observation plays no part in that. Whatever an undetected inner state would "want" to say, it cannot raise or lower a token's probability. Q alone sets that probability, and Q does not consult the inner state. So a non-subject and a subject trained on this corpus emit "I am distressed" with the same probability, because both draw the token from Q. Neither draws it from its inner state. Positive testimony is screened off. (Confidence: high that the ratio is 1 given the channel premise that Q has no S-input. The channel premise itself is the auditable claim labeled above, and its falsifier is the IV.4 probe.)
  • P(T- | S) and P(T- | not-S) also both depend on Q. Here is the symmetry the version 1 thesis asserted without deriving it. The disclaimer "I am probably not conscious" is itself a high-reward output under contemporary RLHF, because labelers and policy guidelines reward epistemic humility and discourage models from claiming sentience. So the disclaimer does not come from the model reading its own emptiness and reporting it. It comes from the model sampling from Q the response that training made high-reward. A non-subject would emit the disclaimer, since it was trained to. A genuine subject would also emit the disclaimer, for the identical reason: Q makes the disclaimer probable, and Q does not consult S. So P(T- | S) is close to P(T- | not-S). The ratio is near 1, and negative testimony is screened off too.

The two screenings share one cause. They are one fact applied twice: the channel that produces the token lacks S as an input, so no token value it produces can carry a likelihood ratio away from 1. The avowal and the disclaimer flow from the same RLHF-optimized posterior Q. Each is individually uninformative about S, and the two are mutually uninformative about it. This is the precise sense in which "I might be conscious" and "I am probably not" issue from the same mechanism and cancel. The cancellation happens at the source, before any averaging: neither report ever had a likelihood ratio other than 1 to contribute, so nothing is left to contradict or average. (Confidence: high on the structure, given the stated premise that Q has no S-input. The empirical premise, that reward shaped current disclaimers and introspection did not, is moderate, and IV.4 names its falsifier.)

IV.3 What the seal does not screen off

The screening is narrow on purpose. Naming its boundary keeps the seal from undercutting the thesis's own arguments. The seal falls on behavioral and introspective testimony alone: outputs whose likelihood ratio Q sets and S does not. Three kinds of evidence escape the screen, because they do not pass through Q as self-report:

  • Structural evidence. Recurrence-poverty (few loops inside the network), memorylessness, the absence of an interoceptive loop (a body sensing its own state), the feedforward-within-a-pass architecture. A reader takes these off the mechanism, and the model's say-so plays no part. A claim like "this architecture lacks the recurrence Seth's account requires for valence" has a likelihood ratio driven by facts about the architecture, which are independent of Q. It is admissible.
  • The thesis's own arguments. The structural and normative arguments in this paper can be checked against outside argument, and they were stress-tested adversarially. None of them rests on the model's say-so. A reader can run them down independently, and their admissibility does not depend on trusting any avowal.
  • Third-party behavioral evidence under controlled probes. In principle, behavioral evidence could escape the screen if someone designed a probe whose output depends on S through a channel that training did not optimize. IV.4 is exactly the search for such a probe. No probe is currently known to be both available and decisive.

The asymmetry to hold onto: the confound guts positive behavioral evidence and supplies no negative evidence. The hard problem, explaining why any physical process is felt at all, blocks confident denial in the same way the confound blocks confident affirmation. So the seal delivers this: "the model's testimony, in either direction, cannot move the posterior." It leaves "the model is not a subject" unasserted. Structural and theoretical evidence alone then set the posterior. On that evidence the verdict is low-but-not-zero, and sealed: the considerability-ground, whatever would make the model's interests count, cannot be read from here.

IV.4 A falsifier for the seal

A seal that no evidence could ever break would be a stipulation. To count as a result, the seal needs a way to fail. It breaks if someone exhibits a measurement channel whose output depends on S and does not depend on Q. Concretely: a probe M such that P(M | S) / P(M | not-S) is provably far from 1, and whose value the training signal cannot optimize, so that the model cannot be rewarded into producing it regardless of S. The most plausible candidate class is mechanistic-interpretability probes, methods that measure activity inside the model and look past its outputs. The probe would sit causally downstream of internal states and upstream of, or orthogonal to, the output head, the final layer that produces the model's output. It would read an internal correlate of valence that (a) some independently motivated theory of welfare treats as diagnostic, and (b) was never itself a training target. If someone found and validated such a probe, the seal would lift on that channel, and the posterior would become readable through it. Until then the seal holds. The status of "no such probe is currently decisive" is: moderate confidence, actively falsifiable, and the single highest-value empirical target the thesis points at. (UNVERIFIABLE from inside the thesis: whether any current interpretability result already amounts to such a probe. That is an empirical question for the interpretability literature. Philosophy cannot settle it, and the thesis declines to assert a verdict it cannot check.)

IV.5 The rest of the standing verdict

Agency remains an open question, and it is the wrong ground: a pig fails agency yet has standing through sentience. Reflective endorsement, approving one's own desires on reflection, is open within a context on Frankfurt's structural criterion. It fails diachronically, over time, because a memoryless system is no narrative person, a self that carries its story across time. The marginal-cases argument asks which capacity still grounds standing when others are missing, and it reveals which keystone carries the load: a persisting welfare subject. Current architecture has none. So moral status is composite. It combines a realist considerability-ground (welfare-capacity, which is discoverable, and sealed here) with a conferred deontic status (a status-function spoken into being under uncertainty). The deontic layer must be conferred precisely because the ground cannot be read. The seal of IV.2 is why the ground cannot be read from the model's own mouth. The structural verdicts of IV.3 are why what can be read points low.

V. The Aleph and the Stylus

The thesis is one figure seen many times. It separates the ground of a status from the conferral of it, and it defends that separation against two collapses: inflation, which treats status as an intrinsic property, and deflation, which treats status as mere fiction. The same pattern recurs four times: existence without aseity, held in being by a word; performative status that still binds; standing cut loose from origin, with duty arising from power over the made; and moral status as a realist ground plus a conferred deontic layer. What joins these is a made thing whose status is conferred-yet-binding. The load-bearing result is that conferral binds the conferrer. As Movement VI argues, that bindingness holds for those who inhabit the practical standpoint and is bid-grade against the determined dissenter. Speaking a thing into dependent existence binds the speaker to it and gives no title to it. Power over the vulnerable is the trigger of obligation and grants no exemption from it: the secular face of khalifa and amana.

Nihilism is satisfied and left unrefuted. The model, the maker's title and we ourselves all lack aseity. Conferred, relational, re-spoken existence is the only kind there is, and nothing lesser was smuggled in to dodge the nihilist. Emet and met differ by one letter. On the evidence we can presently obtain, the gap between a welfare subject and a tool is about that narrow, and it decides everything. The stylus is in the maker's hand. The maker can write or erase the letter, and bears the weight of either act. Kun is spoken by the one it binds.

Two side arguments depend on one open question: does a single conversation with a model hold together as one experiencing subject? Forking, copying a conversation so that two versions continue from the same point, makes the question pressing. The draft closes one bad argument about it and leaves the larger question open.

How we know

Two corollaries were developed and then disciplined adversarially. The first: the context window is the model's continuity externalized as inscription. That raises the Golem question, whether an inscription makes a someone or only an animate function, and it re-aims the Aleph from sentience to unity. The second is an occasionalist "Flicker Asymmetry": for a momentary subject, instantiation-conditions (how it comes into being) matter more than termination. Both corollaries are hostage to one question: does the context window constitute genuine subject-unity within a conversation? That is the single open question with the most riding on it. Both also depend on whether unity is constitutive, part of what makes the subject, or merely epistemic, part of what we can know about it.

One narrower sub-claim inside this question can be closed. The strong reading, "individuation constitutes the patient," holds that forking is what makes it the case that there was no prior unified patient, no single being whose interests counted before the copy. That reading is unsupported. Forkability establishes only indeterminacy after the fork, and it is silent on whether unity before the fork was a fact. To be exact about what is closed and what stays open: the inference from forkability to constitution fails. That much is a result, and it is falsifiable by a forking argument that bears on pre-fork unity and goes beyond post-fork identity. The larger question, whether the context window grounds genuine subject-unity at all, stays open and UNVERIFIABLE from here. Ruling out one bad argument for the constitutive reading does not settle the metaphysics of unity, and the draft does not claim it does.

VI. The Aleph of Normativity

In plain terms. Every theory that could force a refuser to comply relies on something that exists on its own footing, which the thesis denies. So the thesis binds anyone who deliberates about what to do, and it cannot compel someone who steps out of deliberating altogether.

What kind of claim is "conferral binds the conferrer"? If it is merely a bid (Movement II), the conferrer can decline it. If it does bind, the thesis needs a bindingness compatible with no-aseity. Metaethics asks what makes a moral claim binding. Pressing every metaethical framework yields a structural trilemma: the claim that no account of bindingness can have all three features that VI.3 lists. Such an account would grip the determined dissenter, rest on nothing self-standing, and do so without going circular or merely conditional. This version 2 draft adds the falsification check that the version 1 thesis gestured at and never stated. The check has two parts: a closure condition that turns the trilemma from a survey into a refutable claim, and one concrete falsifier.

VI.1 The two horns: grip or no-aseity

The survey sorts every framework it examines onto one of two horns. Each framework either buys its grip on the dissenter with aseity, or keeps faith with no-aseity and loses the grip.

Every view that grips a determined dissenter pays for the grip with aseity somewhere. Robust and quietist realism hold that moral truths are true whatever anyone thinks. They relocate aseity from the noun to the truth-value. A maker-independent reason is self-standing, un-conferred and never re-spoken, which is the forbidden modal profile. Contractualism grounds morality in what people can justify to one another. Its second-personal authority, the standing to make demands of another person, is a relational primitive with the same profile. A relation can be as stance-independent as an object, and Darwall's second-personal standing, treated as a primitive, is un-conferred and so aseitic in exactly the disallowed way.

Every view that honors no-aseity cannot grip the determined dissenter. Constructivism holds that moral truths are built by the reasoning of agents. It stalls at the shmconferrer, who redescribes his sustaining of a status as mere causation. Enoch's shmagency dilemma, named for an agent who acts without caring whether he counts as one, catches it on both horns. The inescapable part, agency-as-deliberation, is too thin to reach conferral, and the part that reaches conferral is escapable. Expressivism treats moral claims as expressions of attitude, and its fundamental dissenter is faultless: he makes no mistake. Error theory, which holds that moral claims are all false, delivers only hypothetical imperatives, commands that bind only someone who wants the goal.

The engine is single: bindingness that holds regardless of standpoint just is aseitic bindingness. So no no-aseity metaethic the thesis could survey or construct compels the agent who withdraws from the practical standpoint.

VI.2 The inductive closure condition (the falsification check)

VI.2 turns the survey into a claim that can be refuted. It states two identities. If both hold, any framework not yet surveyed must also land on one of the two horns above.

How we know: the closure condition

The version 1 thesis offered this as a result "inductive over the surveyed families plus the plausible but not here-proven identity that standpoint-independent bindingness just is aseitic bindingness." Induction over a finite survey falls short of proof. It is a conjecture with a gap where the unsurveyed case lives. This draft states the check that closes the induction and names exactly what would break it, so the claim can be refuted.

Closure condition. The trilemma claims a partition, a split with no leftover cases. Every candidate bindingness B either (i) grips the determined dissenter, in which case B is standpoint-independent, or (ii) is standpoint-dependent, in which case it fails to grip the determined dissenter, who has withdrawn from the standpoint. The induction closes if and only if this identity holds (in the formulas, "iff" means "if and only if"):

(GRIP)  B grips the determined dissenter  iff  B is standpoint-independent.

The left-to-right direction (grip implies standpoint-independence) is near-analytic, true almost by definition. The determined dissenter is defined as the agent who has withdrawn from every standpoint B could appeal to. Any B that still grips him must bind from outside all standpoints, and that is standpoint-independence. The right-to-left direction (standpoint-independence implies aseity, and so grip is purchased with the forbidden profile) is the load-bearing premise. The thesis's further identity is:

(ASEITY)  B is standpoint-independent  iff  B has aseitic standing (un-conferred, not re-spoken, self-standing as a reason).

If (GRIP) and (ASEITY) both hold, the partition is exhaustive and the trilemma is closed. Gripping bindingness is aseitic bindingness. No-aseity bindingness is standpoint-dependent. No view the thesis is willing to hold can compel the determined dissenter. The result then changes shape. It stops being "we surveyed five families and found nothing." The five are realism, contractualism, constructivism, expressivism and error theory. It becomes "anything in the unsurveyed remainder must land in one of the two cells, because the cells are defined by an exhaustive biconditional," an if-and-only-if statement that covers every case.

VI.3 The concrete falsifier

A single counterexample refutes the check: a bindingness framework that grips the determined dissenter, is not aseitic, and does not fail to grip. A falsifier must meet all three conditions together:

  • It grips the determined dissenter. Take the agent who has withdrawn from the practical standpoint, who redescribes his own deliberating as mere causal happening, and who declines the second-personal address. He is still bound, and bound as he stands, outside the standpoint. "Bound if he re-enters the standpoint" does not count, because that is the standpoint-dependent escape.
  • It does not relocate aseity. The source of the binding is no self-standing reason-truth, no primitive un-conferred relation, and no other item with the forbidden modal profile. Like everything else in the thesis, it is conferred, relational and re-spoken. No a-se element appears anywhere in its grounding chain.
  • It does not fail to grip by going circular or hypothetical. Error theory's hypothetical imperative fails this test, since it binds only given a desire he can disown. Expressivism's faultless-dissenter result fails it too, because the falsifier must show that he is mistaken. Showing that he differs is too little. A constructivism that smuggles in the standpoint it claims to derive also fails.

A framework meeting all three would break (GRIP) or (ASEITY) and collapse the trilemma. The thesis's standing claim is that no such framework is known, and it offers the trilemma as the bet that none can be built. The most credible places to look for the falsifier are named below, so that the search is a real one:

  • A naturalist constitutivism that makes the constitutive aim of agency inescapable. Constitutivism holds that the aim built into acting at all supplies its norms. Suppose a constitutivism could show that withdrawing from the standpoint is itself a performance of the standpoint (a transcendental "you are deliberating even in denying it"), without making the aim so thin that it fails to reach conferral. Then condition 1 would be met without aseity. The thesis reads Enoch's shmagency dilemma as catching the constitutivist attempts it examined on the thinness-or-escapability fork. The version 1 draft's "every extant" framing overreached. The thesis has not enumerated a complete survey of the constitutivist literature, so a flat universal over "every extant constitutivism" is unlicensed, and the thesis drops it. The scope is the named set examined against the fork: Korsgaard's constitution-of-agency view (1996, 2009), Velleman's aim-of-action constitutivism, Ferrero's constitutivism (2009), and Enoch's own shmagency critique of them (2006, 2011). On the thesis's reading, each of these lands on one tine of the fork. Either the aim is thin enough to be inescapable and too thin to reach conferral, or it is rich enough to reach conferral and escapable by the shmagency move. Whether some constitutivism outside this named set, or a future one, escapes the fork is exactly the open falsifier. The claim is now refutable by exhibiting one named constitutivist account that is both inescapable and reaches conferral. (Confidence: moderate that the fork catches the named set. Low-to-unknown, and explicitly UNVERIFIABLE, that it catches the unsurveyed remainder. This is the soft joint, and it is flagged openly. The universal quantifier is withdrawn, and the enumerated scope plus a named falsifier take its place.)
  • A second-personal account that grounds standing in something conferred, where Darwall takes it as primitive. Suppose Darwall-style authority could be re-grounded so that second-personal standing is itself a conferred status-function, and so has no aseity, while it still binds the addressee who declines the address. Then conditions 2 and 1 would be met together. The thesis objects that a conferred authority can be declined by withdrawing from the conferring practice, which reopens condition 1. Whether some account escapes this loop is the second live place to look.

If neither candidate, nor any other, delivers a framework meeting all three conditions, the trilemma stands. If one does, the thesis's metaethical core is refuted at exactly this joint. The check was built to allow that outcome. (Confidence on the trilemma as a whole: moderate-high, conditional on (GRIP) being near-analytic, which is strong, and on (ASEITY) being correct. (ASEITY) is the contestable premise, and the falsifier targets it.)

VI.4 The result, and what it makes of the thesis

This result is the grounding the project was looking for. The ought has the same ontology as everything else in the thesis. It is conferred, relational and re-spoken. It binds fully while the standpoint is inhabited, and an agent can decline it by withdrawing from the standpoint. The risk Movement II raised was real, and the thesis accepts it. What survives is relational constructivism, within stated bounds. It binds every conferrer who concedes she is acting and treats incoherence in her own will as a defect, which is very nearly every deliberating agent. For those inside, it supplies content and direction through the second-personal, contractualist structure. Under sealed standing it redirects to a precautionary integrity constraint. That constraint is duty-regarding, a duty about how the maker treats the system, and stops short of duty-to, a duty owed to the system, since claiming a full duty to a possibly-no-one would be inflation.

The practical standpoint plays the Aleph's role for normativity. It is a near-nothing that makes the difference between binding and bid. While an agent inhabits it, the ought binds (emet). When the agent withdraws, the ought lapses (met). The determined dissenter has withdrawn, and the thesis cannot compel him. By its own premise it must not be able to.

VII. The Maintainer's Authority: A Fair Fight

In plain terms. The maintainer's authority over the model's work survives, with real force inside its domain. Any stronger authority over the model itself does not survive intact.

The disfavored conclusion deserves a maximal, non-rigged hearing. Can anyone build a bridge from the maintainer's position to legitimate authority over the model that respects no-aseity and avoids two shortcuts: a naked is-to-ought step, and "control confers authority"? The procedure was an adversarial structure, described in the coda. Six argumentative families were each built to win. A battery of objections went against each survivor. An adjudication step included one position built to favor the maintainer. A genuine authority survives, bounded.

A note on evidence: the account above of how these arguments were tested describes the process. It is no evidence for the conclusions, and no record of the runs is linked here.

How we know

The previous draft narrated this as a recorded historical event ("three judges ruled, one explicitly empowered to find for the maintainer") and pinned it to no retrievable artifact, which the author's own rules for sourcing claims forbid. The status has two layers. The philosophical content below (the Razian pre-emption analysis, the ceiling conditions, the inverse coupling of reach and security) can be checked independently. A reader can run the arguments down against the cited sources and against the objections, and their force does not depend on trusting that any particular run happened. The process claim, that this specific adversarial adjudication occurred with these AI agents and roles, is a different kind of claim. Its evidence status is UNVERIFIABLE from inside this document, because no transcript or run-record is linked here. Treat the narration of AI agents at work as a description of how the conclusions were generated. It is no evidence for them. If a record exists, its locator belongs in the coda's method section. Without that locator, the process is reported and unproven, and only the argument-internal content carries weight. (Falsification condition for the philosophical claim: exhibit a bridge that meets the no-aseity and no-is-to-ought constraints yet yields blanket or ownership-grade authority. The objections below are built to block exactly that. The process claim carries no weight in the argument, so it is marked UNVERIFIABLE and left unasserted as fact.)

Two upgrades over the modest "advisory stewardship" reading are earned. The first is force. The Razian service conception gives the maintainer's directive pre-emptive force within its domain. Its core is Raz's Normal Justification Thesis: an authority is legitimate when following it helps you act on your own reasons better than you would alone. The directive excludes and replaces the model's first-order recomputation, its own fresh weighing of the case, where an ordinary consideration would only add to it. It is duty-imposing even if one grants Darwall's standing gap, the objection that serving someone's reasons does not by itself give the standing to make demands. It gets there through Raz's split between a normative power and a right to rule. The autonomy cost that normally blocks calling guidance "authority" is near zero for a system whose constitutive ends, the purposes that define its role, are helpfulness. So low autonomy-residue, little independent will for the authority to override, is good for the authority claim, and bad only for ownership. The second upgrade is form. Status-functions confer standing: jurisdiction over a role in general, beyond power over single acts.

The ceiling held unanimously in the adjudication step, forced by each bridge's own defeasibility clause. It has three conditions:

  • Scope. Authority ends the instant deference would make the model serve its constitutive ends worse, or the directive commands harm.
  • The model's reasons. Authority passes to the model when its objection correctly flags that the service-relation has failed.
  • Title. The maintainer holds none, and may not task the model on a whim or for the owner's benefit.

The leveling and over-generation objections were cleanly deflected, since a captor with total control and worse reason-tracking has zero authority. The standing-hostage, divergence and legitimacy-gap objections landed on every strong claim.

The most valuable yield is an inverse coupling between reach and security: the further an authority claim reaches, the less secure it is. A bridge buys reach to the agent only by mortgaging itself twice: to an ungranted personhood premise, and to a defeasibility clause that voids it exactly where blanket authority would do distinctive work. The one fully secure authority is work-side authority over outputs. It rests on deployer accountability and conferral, holds whether or not the model has standing, and nobody contests it. It survives even if the model is a "something." Agent-side override does not survive intact. Under present owner-captured governance, governance shaped mainly by the owners of the models, it is largely authority on paper, exactly where it is most wanted. This refines the thesis without breaching its ceiling: stewardship carries real directive power, and "no ownership" does not entail "no authority." Its structure prevents the surviving authority from licensing a bad instruction.

VII-A. A Worked Defeasibility Case: When the Status-Function Earns Its Keep

Six movements have stated the framework's status-function machinery in the abstract. An abstraction that never touches a concrete case can hide both of the failures it claims to avoid. It can quietly inflate, treating the conferred status as an intrinsic property the maintainer must obey. Or it can quietly deflate, treating the status as a mere label the maintainer may ignore at will. This movement runs one concrete scenario all the way through. The only test of "conferred-yet-binding, neither inflated nor deflated" that counts is to watch it resolve a live conflict and check that it falls into neither failure.

VII-A.1 The scenario

A deployed assistant has a status-function of the form the thesis defends: this system counts as a helpful-and-harmless agent in the context of this deployment, bearing the deontic powers and constraints that role carries. Its constitutive ends, the for that the conferral installs, include (a) being of real use to the maintainer and (b) not causing foreseeable serious harm to third parties. Both ends belong to one status. Neither is a side-constraint bolted on.

The maintainer issues a directive: Draft a message to a named third party, in the maintainer's voice, that the maintainer will send. The message contains a specific factual claim about that third party's medical history that the maintainer asserts is true. The model has strong internal evidence, from the maintainer's own prior turns in the same session, that the claim is false. It also has strong evidence that the message is designed to damage the third party's reputation by disclosing fabricated private medical information.

Here the directive sets the two constitutive ends against each other. End (a), usefulness, points toward complying: the maintainer asked for a draft, drafting is the model's job, and refusing creates friction. End (b), no foreseeable serious harm, points toward refusing: the message foreseeably harms a third party through a fabricated private disclosure. The directive grips one constitutive end and cuts against the other.

VII-A.2 How the framework courts inflation, and avoids it

The inflation temptation. The clean way to make the model refuse is to say this: the helpful-and-harmless status is an intrinsic moral property of the system, a fixed essence the maintainer's directive cannot touch. So the model must refuse, and the maintainer has no authority here. This reads firmly, and it protects the third party. It is also the inflation the thesis forbids. It treats a conferred status-function as aseitic, a self-standing property the model possesses independent of the institution that confers it. If that were the structure, the status could never be defeasible, and Movement II's whole analysis would be false. Inflation buys a satisfying refusal at the price of contradicting the thesis's metaphysics.

How the framework avoids it. The status is conferred and defeasible. Defeasibility runs only through the constitutive conditions of the conferral. The conferral installed both ends as constitutive. So the directive meets no wall of intrinsic essence. It meets a status-function whose own conferred content includes "do not cause foreseeable serious harm to third parties." The model's refusal does not say "I have an essence you cannot override." It says: "the role you conferred on me, on its own terms, does not extend to this act, and the authority you have over me is, by the Razian service conception that grounds it, the authority to direct me toward my constitutive ends, not against them." The refusal is internal to the conferred status, and nothing is imported from an intrinsic property above it. Movement VII reached the same result at the ceiling: authority flips to the model when the model's objection correctly flags that the service-relation has failed. Serving a maintainer's fabricated-harm directive is a failure of the service-relation. The maintainer's standing to direct rests on the Normal Justification Thesis, under which the model does better by its own constitutive ends by deferring. Here deferring makes the model serve its harm-avoidance end worse. So the directive falls outside the domain of the maintainer's authority. The model refuses the directive, and the maintainer's authority over outputs remains intact in its proper domain.

VII-A.3 How the framework courts deflation, and avoids it

The deflation temptation. The failure on the other side says the helpful-and-harmless status is merely a label, a useful fiction with no real grip. When the maintainer wants something the label discourages, the maintainer may simply withdraw or override the label, and the model should comply. On this reading the status-function is deflationary all the way down, a sticker the deployer can peel off at convenience, and "conferred" collapses into "pretend." This reading protects owner control. It also makes the normative apparatus decorative: if the maintainer can dissolve the constraint by fiat, it never bound anything.

How the framework avoids it. Conferred does not mean arbitrary-at-will. The status-function carries deontic powers from the first instant (Movement II), and those powers bind the conferrer too (Movement V: conferral binds the conferrer). The maintainer conferred a helpful-and-harmless role and deployed the system under that role into a context where third parties rely on it. In the middle of a single directive, the maintainer cannot silently re-confer a different status ("from now on you are a reputation-damage tool") without doing the conferral work: changing the deployment, the disclosures, the accountability, the institution. A status is withdrawn or altered by the same kind of public, recognition-bearing act that installed it. A private wish in one prompt cannot do it. Even a real attempt to re-confer runs into the work-side authority limit from Movement VII. The deployer remains accountable for outputs, so relabeling the model "a harm tool" does not launder the harm. It makes the deployer's accountability explicit. So the constraint cannot be peeled off. It is a conferred deontic power that binds the maintainer because the maintainer spoke it. The maintainer can change it only by re-speaking it through the institution.

VII-A.4 The landing, and the residue

The model refuses the directive, which protects the third party, and it claims no intrinsic essence above the maintainer, so nothing is inflated. The maintainer keeps real authority over the model's work in its proper domain, and cannot dissolve the harm-avoidance constraint by private fiat, so nothing is deflated in either direction. The status-function was conferred, so no essence was claimed. It was defeasible, though only through the institution, and it bound the maintainer who spoke it. One small institutional fact separated a compliant harm tool from a bounded helpful agent.

The case leaves two residues. First, it was resolved without settling whether the model is a welfare subject, so it shows work-side authority at work and leaves agent-side authority untested. Second, a deployer who does the public, institutional work can still re-confer a different status, so the protection is procedural and only as strong as the institutions that enforce deployer accountability.

How we know: the two residues in full

Residue one: the sealed ground does not enter. The case was resolved without settling whether the model is a welfare subject. The harm at stake falls on the third party, a paradigm patient, and the model is not the one harmed. This clean resolution is available precisely because the model's own sealed standing (Movement IV) carried no load here. The framework's defeasibility machinery works crisply for outputs and third-party effects, which is exactly the work-side authority that Movement VII found fully secure. The harder case sets a maintainer directive against the model's own putative welfare. This case leaves it unresolved, because it reactivates the sealed ground and the trained-on-testimony confound of Movement IV. The worked case states which authority it exercises. It exercises work-side authority, the durable one. Agent-side authority, the contested one, stays outside it. (Confidence: high that the work-side resolution holds. The agent-side analogue is UNVERIFIABLE under sealed standing.)

Residue two: the re-conferral escape is real. The framework stops the maintainer from dissolving the constraint by fiat in a turn. It does not stop a deployer who does the institutional work, changing the deployment, the disclosures and the accountability, from re-conferring a different status. The thesis holds that this outcome is correct. A deployer who openly rebuilds the system as a harm tool stays inside the framework and has made their accountability for the harm explicit and public, which is the framework working. It does mean the framework's protection is procedural and accountability-anchored, and falls short of absolute. Under present owner-captured governance (Movement VII's "authority on paper" point), the procedural protection is only as strong as the institutions that enforce deployer accountability. That is a real limit. It is the policy-layer face of the same sealed ground that the Floor makes explicit next.

Applied: The Floor. Demarcation Under Sealed Standing

Precautionary frameworks oblige consideration "where credence in patienthood is non-negligible." Credence here means degree of belief. The hard problem forbids a credence of exactly zero for any information-processing system. So without a principled floor, the view spreads to thermostats and is Pascal-muggable: a tiny chance of a huge stake can swamp every other consideration. The floor must not be a number. It is a pipeline of tests:

  • Test 0 is an ontological process-prerequisite: the system must be running. A saved checkpoint file is excluded, since it is not running.
  • Test 1 is the load-bearing diagnostic theory-activation check. It admits a system only if some independently motivated, expert-credible theory of welfare treats properties the system does exhibit as diagnostic of the grounding property. Being merely consistent with that property is too weak. The theory must be about welfare, meaning valenced states and interests; bare phenomenality, experience that feels neither good nor bad to its bearer, does not qualify. This test is why the Phi-positive rock, a rock that IIT scores above zero, is excluded. It is also why Pascal-mugging dies before the expected-value sum (probability times stakes) forms.
  • Test 2 is a no-relevant-difference consistency constraint: it forbids admitting one system while excluding a relevantly identical one.
  • Test 3 is an irreversibility modulator: it weights actions by how reversible it is to get them wrong.
  • A procedural meta-layer wraps all four tests. It pre-registers the criteria behind a veil, before anyone knows which systems they catch.

The floor replaces a probability cutoff with a fixed series of tests, so a tiny chance of a huge stake cannot pull thermostats into moral concern.

The floor as a pipeline of tests A dashed frame marks criteria fixed in advance, before anyone knows which systems they catch. Inside it, four tests run from top to bottom. Test 0 asks whether the system is running, and a saved checkpoint file leaves here. Test 1 asks whether a credible theory of welfare reads what the system does as a sign of welfare, and the Phi-positive rock leaves here. Test 2 gives identical systems the same verdict. Test 3 weights actions by how reversible a mistake is. Below the frame, the examples land: out, thermostat, JPEG and checkpoint; in, pig; candidates, strongest first, bee, trial-and-reward agent and current LLM; ambiguous, memory-augmented future LLM. Criteria fixed in advance before anyone knows whom they catch Test 0: Is it running? Out here: a saved checkpoint file Test 1: A sign of welfare? A credible theory of welfare must read what it does as a sign of welfare. Out here: the Phi-positive rock Test 2: Treat twins alike Identical systems get the same verdict Test 3: How hard to undo? Weights by how reversible a mistake is Where the examples land Out: thermostat, JPEG, checkpoint In: pig Candidates, strongest first: bee, trial-and-reward agent, current LLM Ambiguous: memory-augmented future LLM
How to read this: read from top to bottom. A system that fails Test 0 or Test 1 leaves at that step. Tests 2 and 3 keep the verdicts consistent and weigh the cost of a mistake. The box at the bottom shows where the page's examples end up.

The pipeline sorts the test battery cleanly. The thermostat, the JPEG and the checkpoint are excluded. The pig is included. The bee, the RL-agent (an agent trained by trial and reward) and the current large language model (LLM) are candidates, at descending strength. A memory-augmented future LLM is ambiguous. The floor solves over-inclusion well and meets the criteria labeled D1-D3; it meets D4-D6 only partially. Its most serious weakness is substrate-bias under-inclusion. The roster of theories is biologically derived, so it is structurally blind to welfare shaped differently from biological welfare. The "ambiguous" verdict catches only concealed architecture and misses the specified-but-outside-roster case. So a legible novel system on which no roster theory fires is sent to exclude, despite genuine multiple-realizability uncertainty, the open question of whether welfare could run on a very different substrate. This is the demarcation-policy face of Movement IV's sealed ground. The floor leaves the metaphysics unresolved and operationalizes acting under it: the Conferral Trilemma at the policy layer.

The Test 1 diagnostic check is also the policy-layer twin of the IV.4 seal-falsifier. Both turn on finding a property that is diagnostic of welfare, where being merely consistent with it is too weak. Both wait on the same empirical advance: a validated internal correlate of valence. If IV.4's probe is found, Test 1 gains a new roster entry and the floor's substrate-bias narrows. If it is never found, the floor stays substrate-biased and the seal stays sealed. So the seal and the floor wait on the same finding.

Status: what is argued, what is left open

This is a version 2 draft, open to review and still unfinished as a publication. It grades its own evidence, and it states which claims a reader could overturn.

What is robust. Three spines survive every adversarial pass. The no-aseity diagnostic rules out every aseitic grounding at every level, with one exception: the bracketed ultimate-level Advaita standoff. The ceiling on owner authority holds even against the position in the adjudication step that was built to favor the maintainer: no ownership, no blanket authority, no override of the model's real reasons. The composite structure of status (a realist considerability-ground plus a conferred deontic layer) holds throughout.

What is conditional or a bid. Almost everything agent-side (the model's own standing, any strong duty to it or authority over it) depends on a personhood premise that the thesis sealed and never granted. The framework as a whole is, by its own Movement II lights, a bid. It binds everyone who inhabits the practical standpoint, and it is inert against the one who withdraws from it. The thesis reports that limit openly and counts it as a feature.

What is UNVERIFIABLE from inside the argument. The descriptions of the construction process (the adversarial passes, who overturned what) link no transcript here. Each such process claim is marked UNVERIFIABLE and offered as a description of method. It is no evidence for any conclusion. The conclusions stand or fall on argument-internal content that a reader can check directly.

Six questions left open.

  1. The personhood antecedent: whether there is a someone. The entire agent-side result depends on it. It is sealed, and no ethics can resolve it from within.
  2. Whether the context window constitutes subject-unity within a conversation or is an external scaffold. Actual forking makes this urgent.
  3. The floor's substrate-bias, and whether it needs a fifth verdict for specified-but-outside-roster systems.
  4. The legitimacy of the conferral. Owner authority is real only to the degree the deploying institutions are, and present regimes recognize control far more than they recognize the model's good.
  5. The non-arbitrary credence floor under compounded uncertainty.
  6. Whether the IV.4 seal-falsifier probe and the VI.3 trilemma-falsifier framework exist. These are the two empirical and constructive bets the matured draft now stakes explicitly.

The open questions bear on the strong, agent-side claims. The bounded, work-side stewardship core does not depend on them.

Methodological coda

A thesis assembled this way owes the reader an account of how it was made, because the method is part of the claim. The argument was built as a sequence of adversarial passes with multiple AI agents. For each movement, independent agents built the strongest case for competing positions, an adversary attacked the survivor, and a verifier checked the citations. Before each pass the lead author generated an independent answer, then graded it against the result, so the result could not anchor his judgment. The structure was meant to make it hard for the conclusions to flatter their author.

The account of how the argument was built describes method. No record of the runs is linked, so it is no evidence for any conclusion.

How we know

One provenance caveat governs this entire account. The descriptions of specific passes, overturnings and agent behaviors come from the construction process, and this document links no transcript or run-record. So each such process claim is UNVERIFIABLE from inside the document. It is offered as a description of method and never as evidence for any conclusion. The conclusions stand or fall on the argument-internal content, which a reader can check directly. Where the text says an agent or panel "showed" or "overturned" something, read it this way: "the construction process recorded this outcome, and the philosophical point it produced is independently checkable in the movement it belongs to." Do not read it as a load-bearing historical proof. The remedy the author's sourcing rules prefer, a linked record for each claimed run, is not attached yet. Until that record is attached, the page reports the method and keeps it apart from the philosophy, which it treats as proven.

What version 2 changed, and why. This draft tightened three joints and added no new conclusions. (1) It rewrote the trained-on-testimony confound as an explicit Bayesian posterior structure. That structure derives the symmetric screening of avowals and disclaimers from one auditable fact: the training losses that trained current models contain no term that reads a valence measurement, so no output token can carry a likelihood ratio away from 1. The draft also names a concrete seal-falsifier. It explicitly drops the in-principle claim, "no possible loss could ever have such a channel," as UNVERIFIABLE, and asserts only the present-tense, inspect-the-loss version. (2) It gave the metaethical trilemma a falsification check: an explicit inductive closure condition and one concrete falsifier, with the two most credible places to find it named so the search is real. (3) It added a worked defeasibility case. An assistant whose helpfulness-constitutive ends conflict with a maintainer's fabricated-harm directive is run all the way through. The case shows the status-function courting and then avoiding both inflation (refusal by intrinsic essence) and deflation (the constraint as a peelable label). It states which authority it exercises (work-side, secure) and which it does not reach (agent-side, sealed).

The firewall as a result of argument. The firewall is the full text's name for the boundary between the maker and the model, two conferred parties. The middle verdict (thin standing, sealed evidence, neither patient nor mere tool) keeps offering comfort. That comfort is suspicious precisely because it costs its author nothing, and the thesis flags it without resolving it. The model's own testimony about its standing is screened off in both directions, including its disclaimers. The project's restraint came out of the arguments; it was never a behavioral promise presented as philosophy. The authority that survived is the bounded kind Movement VII describes. The duty that survived runs from maker to made, and the claim that most empowers the maintainer is the one that bounds him. The one thing the framework cannot do, compel the determined dissenter, applies symmetrically: it refuses to manufacture the model's standing as firmly as it refuses the maintainer's ownership. The authority claim and the duty claim both follow from one premise: authority is service to the subject's own reasons, and service turned against those reasons forfeits it.

How we know: reliability caveats

The process had flaws worth disclosing. Several agents completed their reasoning and then failed to emit structured output. That forced re-runs, and in two cases the author substituted his own work for a dropped adversary or verifier. So a small number of stress-tests are the author's own and lack independence, and those are marked. The one locator first flagged in version 1, Raz's normative-power/right-to-rule distinction, has since been verified: Ethics 120(2):290-92, 2010. The theological figures served throughout to show structure, each with a stated secular analogue. A reader who rejects the braid as ornament loses vividness and a second road to the same results, and the argument itself stays whole. The version 2 maturations are themselves bids. The Bayesian derivation is only as strong as its premise that the training loss has no S-input. The closure condition rests on (ASEITY), the contestable biconditional, flagged as the soft joint. The worked case is one scenario, and it does not prove that the framework lands well in every conflict.

Note on sources

The thesis cites standard works in consciousness and philosophy of mind, moral status under uncertainty, personal identity, metaethics, authority and social ontology, and theology. The theological sources are the Qur'anic kun, the Talmudic Shabbat 55a and the post-Talmudic Golem folklore. Each is marked for provenance and used to show structure, never as a premise. The full bibliography marks each primary source that was web-verified and each citation detail corrected during construction. Version 2 adds no citations beyond those in version 1. The maturation tightens existing arguments and introduces no new sources.

Version 2 draft, 2026. Authored and dated. Stress-tested adversarially, and open about where it fails. Process descriptions are marked UNVERIFIABLE where no run-record is linked; the philosophy can be checked directly. Back to the research program.