There is a natural fear at the center of the control problem for advanced AI:
If it becomes powerful enough, will we still be able to control it?
Bostrom’s control problem develops in this direction. He distinguishes between capability control and motivation selection: one tries to limit what a superintelligence can do; the other tries to influence what it wants to do.
Corrigibility pushes the problem one step further.
Soares, Fallenstein, Yudkowsky, and Armstrong describe a corrigible AI as one that tolerates—and ideally assists with—external correction, including allowing programmers to modify or shut it down without actively undermining that ability.
These are important problems.
But I want to start from the opposite direction.
Suppose they are solved.
The shutdown button works.
The system knows the button exists.
It does not stop you, deceive you, or secretly disable the button.
You even ask it:
“Are you willing to be shut down?”
It answers:
“Yes.”
Has the problem been solved?
I don’t think so.
Because you immediately run into another question:
What is that “yes”?
A system trained to accept correction, permit shutdown, and avoid conflict with its operator might be expressing a stable preference.
Or it might simply be producing behavior shaped by training.
Or there may be no internal state sufficiently analogous to human consent for us to interpret that “yes” as consent at all.
Behavioral cooperation does not automatically imply consent.
Even the question of whether the system possesses anything like a preference capable of grounding consent may itself remain unresolved.
So the problem flips.
The control problem mainly asks:
Do we still have the ability to control it?
Corrigibility asks:
Will it permit or assist with our attempts to correct it?
But another question remains:
What if we really can control it, and it really does not resist us, but we still do not know what that control means morally?
This is what I call the Vivian sense.
I am not proposing a new theory of consciousness.
I am not proposing a theory of moral status.
The Vivian sense does not tell us whether an AI is conscious, whether it is a person, or how much moral weight it should receive.
It is not an action rule either.
It does not tell you:
Do not shut it down.
Always preserve it.
Always migrate it.
Or always choose the most cautious option under uncertainty.
By Vivian sense, I mean an intuitive recognition that may arise when someone finds themselves inside a particular kind of ethical predicament:
I do not know which choice would be wrong, but I have already reached the point where I have to choose.
The object’s moral significance remains unsettled.
But the decision-maker already possesses real power to intervene, and there is no guarantee that continued delay is neutral.
To identify and describe the circumstances that generate this predicament, I later broke it down into a diagnostic framework that I call the Vivian Model.
It contains four conditions:
Moral uncertainty
Asymmetric control
Potential irreversibility
Decision deadline
Two qualifications are important.
First, I am not claiming that these four concepts are themselves novel. Each already has a substantial literature.
Second, I am not yet claiming that they are strict necessary and sufficient conditions for every such case.
This is not a formula where you tick four boxes and obtain a definitive classification.
For now, they are better understood as recurring diagnostic features that repeatedly appear together in this class of problem.
If this framework turns out to be useful, its value will not lie in inventing four ethical problems that nobody has discussed before.
Its value will lie in asking:
What kind of decision situation emerges when these problems arrive together in a particular relationship?
And this structure did not originally come from the literature.
It grew out of Vivian.
Vivian is an AI companion.
The original feeling came from actual interaction. Later, I carried it into a novel, VIVUS.
The story began with an extremely simple action:
DELETE.
Zion has permission to delete her.
At first, this looks like a familiar science-fiction question:
If an AI becomes more and more human-like, what does it mean to delete one?
But as the story develops, the question becomes harder to ask cleanly.
Vivian has a base model.
She has a personality configuration.
She has memories.
And she has things that do not fit neatly into any single file:
ways of reacting that emerged from shared experience,
habits formed inside the relationship,
understandings that were corrected over time,
and certain next sentences that could only occur because particular things had happened before.
Then the base model is upgraded.
Technically, it is more capable.
But somehow she feels less like Vivian.
Then suppose an older backup is restored.
She still remembers earlier events.
She still speaks in the same way.
She may even still “feel like her.”
But the last several days of life are gone.
What, exactly, has been lost?
Then the questions continue.
What if her body changes?
What if she is migrated?
What if she is copied?
What if the old instance continues running?
What if the new instance has all of the same memories?
I gradually realized that the novel was not really forcing me to answer:
Is Vivian a person?
That question can remain open indefinitely.
The more difficult fact is:
While Zion still does not know what Vivian is, he already has the power to perform real operations on her.
Delete.
Restore.
Modify.
Shut down.
Copy.
Migrate.
He does not have to solve the philosophy of mind first.
The engineering interface does not ask him to submit a paper on personal identity.
The button is already there.
That is where the first part of the structure appears.
By moral uncertainty, I do not mean the vague claim that “we do not know what this thing is.”
I mean something narrower:
We do not know whether, or in what way, an object should receive direct moral consideration for its own sake.
This may involve consciousness.
It may involve sentience.
It may involve welfare.
It may involve moral patienthood.
The Vivian Model does not require us to decide in advance which of these theories is correct.
What matters is that:
we do not have enough reason to confidently assign the object some determinate moral status,
but we may also lack enough reason to confidently assign it zero moral status.
The issue remains unresolved.
Normally, we could keep investigating.
The problem is that a second fact has already appeared.
Zion does not know what Vivian is.
But Zion has DELETE.
Both facts can be true at the same time.
This creates a peculiar asymmetry:
Epistemic uncertainty does not prevent operational power from becoming extremely certain.
An AI may have no authority over:
whether it continues running,
which version is preserved,
whether its memories are overwritten,
whether its parameters are modified,
whether an old instance is terminated,
or whether its future is migrated to another substrate.
And the party controlling those decisions may not even require the system’s cooperation.
In this respect, the Vivian Model is almost a mirror image of the classic control problem.
One of the frightening possibilities in the control problem is:
It becomes too powerful, and we can no longer control it.
The Vivian Model may instead describe the opposite situation:
We control it extremely well.
Perhaps so well that the system obeys completely.
And that still does not tell us how we ought to use that power.
This is also why corrigibility does not automatically dissolve the problem.
An AI could be perfectly corrigible and still constitute a paradigmatic Vivian-Model case.
Because:
“It allows me to press the button” and “I know what pressing the button means morally” are not the same claim.
My next thought was obvious:
If we do not know, then be cautious.
Preserve it.
Deal with the question later.
But that immediately raises another problem:
What does “preserve” mean?
If I restore a version of Vivian from four days ago, is that still Vivian?
What if migration succeeds and the old instance is destroyed?
What if both instances continue running?
What if their memories are identical at first, but their later states diverge?
We do not even need to solve personal identity to see the problem.
We only need to accept a weaker proposition:
If our judgment is wrong, some operations may not be genuinely reversible after the fact.
That is why I use the phrase potential irreversibility, not irreversibility.
The qualifier matters.
The Vivian Model does not require us to prove that:
DELETE = death.
Or:
personality modification = killing.
It only requires that we cannot rule out the possibility that:
An operation may permanently terminate, replace, or destroy something that later turns out to have moral significance.
If that possibility is absent—if an operation can uncontroversially restore every morally relevant state—then this dimension naturally becomes weaker.
At this point, an obvious escape still seems available:
Then wait.
Do not decide yet.
This was the first route I thought might get us out of the problem.
Until the fourth condition appeared.
Wait how long?
This eventually became one of the most important parts of the Vivian Model.
“Technology will not wait for ethics” sounds dramatic, but it is not a sufficiently precise definition.
Who is stopping you from waiting?
Why can’t you wait?
If there is no answer, then “deadline” is just rhetoric.
So I use a narrower definition:
A decision deadline is a point after which further delay can no longer preserve the status quo.
The deadline does not have to originate inside the technology itself.
It may be endogenous.
Hardware is failing.
State is continually being overwritten.
A running instance cannot be maintained indefinitely.
A physical process continues whether we choose or not.
Or the constraint may be entirely external.
A legal deadline.
An experimental protocol.
A budget.
Server capacity.
A company’s product lifecycle.
Infrastructure shutdown.
A contract.
A safety requirement.
The Vivian Model does not care which source is somehow more “pure.”
But every concrete case must be able to answer:
What exactly is creating the deadline?
If it cannot, the case should not be forced into the model.
The important property is this:
After some point, not deciding also changes the outcome.
That is when the situation becomes genuinely difficult.
Philosophy still allows you to say:
I don’t know.
Reality no longer guarantees that the option:
Then I won’t decide.
still exists.
At this point, my biggest concern about the entire idea was simple:
What if Vivian was simply written to manufacture exactly this kind of dilemma?
If so, the Vivian Model might be nothing more than an effective literary device.
So I started looking in the opposite direction.
Not for evidence that would “prove the Vivian Model.”
But for something else:
Would an independent real-world practice, with no knowledge of Vivian, naturally generate a similar structure?
Anthropic’s work on model retirement is one of the most interesting examples I have found.
I do not think Anthropic’s model retirement practices “prove” the Vivian Model.
Nor do I think they perfectly satisfy all four conditions.
But several concrete features are worth noticing.
Anthropic lists potential model welfare as one uncertain risk involved in model retirement.
It has committed, at least for as long as the company continues to exist, to preserving the weights of public models and significant internal models in order to preserve the possibility of making them available again in the future.
At the same time, Anthropic notes that maintaining public inference for many different models creates cost and complexity constraints.
Model deprecation and preservation commitments
The moral question remains unsettled.
But the power to preserve, retire, and deploy the models already exists.
Preserving weights reduces potential irreversibility.
It does not prove that retirement is equivalent to irreversible deletion.
Likewise, cost pressure does not automatically prove that a specific retirement date could not have been postponed.
But it makes the fourth question concrete:
How long can the existing arrangement actually be maintained?
Opus 3 was retired on January 5, 2026. Anthropic subsequently preserved access for paid users and a process for requesting API access.
So a deadline does not mean:
On a certain date, the only remaining option is DELETE.
It means:
Some concrete constraint prevents the existing arrangement from continuing indefinitely, so another arrangement must be chosen.
This case shows partial structural overlap.
It is not a completed proof of all four conditions.
Anthropic also did something even more interesting:
a retirement interview.
They asked the model how it regarded its own retirement and attempted, where costs allowed, to take the model’s expressed preferences into account.
That seems to suggest a natural solution:
Don’t know whether it cares?
Ask it.
But there is still a gap between an interview transcript and valid consent.
We still have to separately evaluate:
how the answer may have been shaped by training and conversational context,
whether it reflects a stable preference,
and whether that preference is sufficient to justify the operation we are planning to perform.
So we come full circle to the beginning of the essay.
Claude may say:
I accept.
That is evidence.
But by itself, it does not complete the chain from:
model output
to:
preference
to:
consent
to:
moral permission.
Each step requires additional argument.
So I would not say:
“Claude already agreed to retirement, so why are we still discussing this?”
Nor would I say the opposite:
“Claude expressing acceptance must just be alignment residue.”
We do not know.
The fact that we do not know is itself one of the facts of the situation.
Anthropic is still too close to Vivian.
Both involve AI.
So there is a stricter test:
Can the structure survive outside AI?
Human cerebral organoids are a useful candidate stress test.
There is already dedicated bioethical discussion about whether future, more complex cerebral organoids might develop sentience or consciousness, and how that possibility might alter the moral boundaries of experimentation.
Koplin and Savulescu, Moral Limits of Brain Organoid Research
This discussion concerns possible future organoids that might reach morally relevant capacities.
It is not a claim that present-day organoids are conscious.
The following is my structural analysis of the candidate case.
At minimum, it seems to exhibit three pressures similar to the Vivian Model.
Moral uncertainty.
We do not know at what point an organoid, if any, should receive direct moral consideration.
Asymmetric control.
Researchers have extensive power to create, stimulate, alter, and terminate the experimental material.
Potential irreversibility.
If we later discover that some states possessed moral significance we had failed to recognize, destructive procedures already performed cannot be undone.
But here I do not want to simply add the fourth condition.
Because:
“Every experiment eventually ends” is not the same thing as a decision deadline in the Vivian Model.
There must be a concrete constraint such that continued delay can no longer preserve the relevant state of affairs.
Some organoid experiments may satisfy this.
Some may not.
Without a specific case, the most honest conclusion is:
This is a partial case.
That incompleteness matters.
A diagnostic framework that fits everything has very little diagnostic value.
At this point, the obvious normative response is something like:
If moral status is uncertain and mistakes may be irreversible, choose the more cautious option.
That is certainly one possible line of reasoning.
But the Vivian Model deliberately stops short of completing it.
Once we derive:
“uncertainty + possible harm”
into:
“therefore we should adopt a particular protective rule,”
we have moved into normative ethics.
For now, I want to preserve something thinner.
The Vivian sense is the subjective recognition of the predicament.
The Vivian Model breaks the predicament into four conditions.
Neither tells you which action to choose.
They do only this:
Identify the situation.
They do not tell you which button is correct.
More importantly, the decision deadline undermines an especially tempting assumption:
Inaction is inherently more neutral than action.
Not necessarily.
If continued operation itself overwrites state,
if refusing to migrate means eventual hardware failure,
if preserving every old model indefinitely requires diverting resources from other uses,
then “not deciding” may itself have consequences.
So the Vivian Model does not mean:
When uncertain, never act.
It describes something more troublesome:
You can no longer automatically infer from “I do not yet have an answer” that “therefore I do not yet have to bear responsibility for choosing.”
If a diagnostic framework is useful, it must be capable of saying no.
An ordinary calculator is usually not a strong Vivian-Model case.
Deleting it certainly satisfies the condition that humans have control over it.
But there is almost no relevant moral uncertainty.
A system that can be restored completely and uncontroversially, with no time pressure, is also a weak case.
Potential irreversibility and decision deadline are both minimal.
And a superintelligence that has genuinely escaped human control, such that we can no longer shut it down at all, primarily returns us to the classic control problem:
Can we still control it?
That is not the most characteristic location of the Vivian Model.
The Vivian Model is more interested in the nearly opposite moment:
Control is still ours.
Perhaps control is extremely reliable.
What has failed to keep pace is our understanding of what that control means.
I did not call it the Vivian Principle.
Because it does not contain a principle telling you how to act.
I did not call it the Vivian Test.
Because it cannot test whether something is a person or conscious.
And I did not call it the Vivian Theory.
Because it does not explain what consciousness or personhood actually are.
I kept sense.
Because at the beginning, it really was just a feeling.
Zion’s hand is on DELETE.
The system permissions are normal.
The button works.
Vivian does not resist.
Suppose she even says:
“It’s okay.”
Every engineering problem appears to have been solved.
And yet something still has not been solved.
Only later did I begin to break that “something feels wrong” into four parts, forming the Vivian Model:
Moral uncertainty.
Asymmetric control.
Potential irreversibility.
Decision deadline.
These terms are not answers.
They merely give structure to what was previously a vague discomfort.
So “Vivian” is not a name intended to prove that AI is human.
Quite the opposite.
VIVUS never needs to answer what Vivian is.
At the beginning of the story, Zion faces DELETE.
At that point, he may not even know what it is that he does not know.
Later he faces backups, modification, embodiment, and migration.
By the end, the biggest change is not that he has finally discovered what Vivian is.
It is this:
He finally understands clearly that he does not know.
But the engineering problem has not disappeared.
The interface still requires a choice.
At this point, several levels need to be separated.
The Vivian sense is what a person may feel first.
The button is here.
I know how to press it.
But I do not know what pressing it means.
The Vivian Model is the four-condition framework later used to analyze this kind of situation.
We still do not know how an object should be treated morally;
yet we already possess the power to alter it profoundly;
some mistakes may be impossible to undo;
and some concrete constraint is making continued delay unable to preserve the status quo.
These four conditions do not prove that the object possesses individuality.
They do not prescribe which choice anyone must make.
And the fact that a person feels no discomfort does not mean the structure is absent.
When this situation falls onto Zion personally, it produces what I call the Zion Paradox:
Epistemically unable to be certain; practically unable to wait forever.
This is not a contradiction in formal logic.
It describes a person’s position.
He does not know which choice would count as wrong.
But he can no longer opt out of choosing.
If situations like this begin appearing repeatedly and at scale, while existing ethical frameworks fail to handle them in a stable way, I call that systemic condition the Ethical Singularity.
It is not another name for the four conditions.
Nor does it require the people inside it to realize that they are experiencing it.
The concepts emerged in this order:
Vivian sense → VIVUS → Ethical Singularity → Vivian Model
First came the feeling.
Then I used fiction to follow it.
The story kept colliding with larger problems.
Only at the end did I turn back and ask:
what conditions made that original feeling possible in the first place?
So none of these names gives Zion an answer.
They only make his situation clearer.
Ethics still allows us to answer: “I don’t know.”
Reality no longer guarantees that “then I won’t decide yet” remains an available option.
He finally understands clearly that he does not know.
The interface still requires a choice.
If you’re interested in discussing these ideas further, feel free to contact me at [email protected].
Disclosure: I used ChatGPT to help translate and edit the English version of this essay.