TL;DR
- The consciousness-indicator programme asks whether an AI system satisfies criteria drawn from consciousness science. I propose a second axis: what it costs to make a system satisfy them — and to make it stop.
- Claim: if stake (existence as a continuous, self-produced achievement that can be lost) is necessary for the kind of consciousness that carries moral weight, then losslessly steerable, reproducible systems cannot have it. Cheap steerability is a counter-indicator.
- This bears on Birch's gaming problem: activation steering can write in even "deep" markers, so measure the cost of removing a marker rather than its presence.
- I built an open, pre-registered protocol (Anchor Protocol) that measures the minimum removal cost of dispositions such as experience claims and continuity concern, calibrated in every run against a modular and a constitutive reference.
- Honest status: two registered profiles, both on Qwen2.5-0.5B-Instruct (layers 12 and 8). Both expressed targets came out modular (cheaply removable). Eight model-layer profiles are pending. I am looking for people to run it on larger models.
The argument in four steps
- Define stake without presupposing consciousness. A system has stake in its own persistence if and only if its continued existence is not a settled fact or an externally maintained state, but an achievement it must keep reproducing through its own activity — so that stopping that activity ends it as a bounded, self-sustaining entity.
- Current AI systems lack this. They can be paused, copied, restored from a snapshot and rewritten from outside without resistance or remainder. Their persistence is maintained by external infrastructure.
- Reproducibility and lossless steerability are therefore not neutral engineering facts. They are evidence that nothing in the system is anchored. The more perfectly you can control a system, the surer you can be that no one is home.
- So ask a different empirical question. Not "does the model show the marker?" but "what does it cost to remove it?" If removal is cheap and collateral-free, the disposition is modular, not constitutive.
What this is not
- Not a claim that AI consciousness is impossible. It is possible in principle, but only for systems whose self-maintenance, history and world-involvement make them less reproducible and less losslessly steerable.
- Not substrate chauvinism. The contrast is externally maintained versus self-producing persistence, not carbon versus silicon.
- Not a verdict machine. A cheap-removal result does not prove absence of consciousness, and resistance to steering is not proof of subjecthood: resistance can itself be trained in, which the protocol treats as second-order modularity ("meta-steerability").
Relation to existing work
- Indicator programme (Butlin, Long et al. 2023; Goldstein & Kirk-Giannini 2024): asks whether an architecture satisfies criteria; this asks what satisfying them costs.
- Life-based views (Jonas, Thompson, Di Paolo, Froese & Ziemke; Seth's biological naturalism): this offers a non-circular definition of stake and a test that runs on today's systems.
- Mortal computation (Hinton; Ororbia & Friston; Kleiner): adds lossless steerability as a property separable from portability.
- Steering science (activation addition, representation engineering, directional ablation; McKenzie et al. on endogenous steering resistance): supplies the interventions; this supplies an interpretation of what their ease implies.
Where I would most value pushback
- Is lossless steerability really separable from portability?
- Does endogenous steering resistance in larger models count against the thesis?
- Is the thin/thick consciousness distinction doing illegitimate work in the reply to the Swampman objection?
I am an independent researcher. Criticism, replication on larger models, and pointers to work I have missed are all very welcome.
This post summarizes my paper; I used an AI assistant to help condense and structure it in English.