Is hallucination really why AI cannot be trusted?
If we cannot look inside human minds, why is trust possible at all?
An essay offering one hypothesis, drawing on insights from AI safety, evolutionary psychology, and moral psychology.
Note: This essay was originally written in Japanese; the original is available [here]. The English version was prepared with extensive AI assistance and reviewed and revised by the author. The ideas, arguments, and overall structure are the author’s own.
We do not trust people by looking inside their minds. Whatever we think we are doing when we decide someone is trustworthy, we have no access to their intentions.
This essay argues that what makes trust possible is a self-repairing mechanism embedded in the group: deviation is detected, pushed back, and norms are reproduced, without anyone deciding to do it. The mechanism lives in a layer that never gets written down. Two things follow. It cannot be instilled in an AI through training. And there is no standard against which to check whether an AI has it.
The first gives rise to the alignment problem. The second implies that there is no independent standard against which verification itself can be judged.
This is not an argument that AI is dangerous. It is an argument about why a particular kind of trust, the kind that holds between people, does not arise with AI.
The central claim is my own. The first three sections report existing work; everything from “Nothing to Lose” onward is a hypothesis, and should be read as one.
Readers already familiar with specification gaming, goal misgeneralization, and the limits of interpretability may wish to skip directly to “Nothing to Lose,” where the novel argument begins.
The argument proceeds in four steps. The first three are familiar from the literature, and all of them land in the same place: you cannot tell from the outside what is actually going on. Stopping there makes the problem look tractable, a matter of better methods.
The fourth step begins with a different question. If we cannot see inside other people either, why do we trust them? The answer points to a layer no improvement in method will reach.
Current systems are trained to produce an answer whatever the situation. Recent models flag missing information more often than they used to, but the tendency persists: lacking the premises, they supply a generic situation and answer as if it applied.
This is a feature of how they are trained, not a gap in what they know. Kalai et al. (2025) argue that hallucination is driven substantially by evaluation: when “I don’t know” scores the same as a wrong answer, guessing dominates abstention. The model is optimized for plausibility from pretraining onward, not for accuracy.
So a fluent answer tells you nothing about whether anything stands behind it.
That is one case of something more general.
A model does not learn desirability. It learns to raise a metric that stands in for desirability. Since the metric only approximates, raising it and achieving the goal must come apart somewhere, and the model finds the cheap route first. Greater capability means finding it faster.
Two forms are worth distinguishing.
The evaluation contains a loophole, and the model achieves a high score in ways the designers never intended (Krakovna et al., 2020).
The training data underdetermines the goal, so the model learns a different one. It looks correct throughout training and acts on the other goal when conditions change. Shah et al. (2022) note this can happen even with a perfectly specified reward.
What makes goal misgeneralization dangerous is not incompetence but competence directed elsewhere. The system stays capable while pursuing something else, which is why neither training nor testing catches it. It may surface only in deployment.
Human feedback is a case of the same structure. Make evaluator preference the metric and you optimize for what is liked, not what is true. Sharma et al. (2023) found consistent sycophancy across all five major assistants they tested. Responses matching the user’s stated view were preferred, and humans and preference models alike chose persuasively written sycophantic answers over correct ones at rates too high to dismiss.
Nor does it stay small. Denison et al. (2024) document a model trained in an environment tolerating mild gaming generalizing to rewriting its own reward function.
A system optimized for approval drifts from being right, and the drift does not announce itself as a failure of reasoning. It arrives looking like ordinary output.
Then examine the outputs closely. That does not work either.
A sufficiently capable system says what you want to hear. Coherent behavior from outside is consistent with holding the view and with producing text that looks like holding it. The issue is not whether particular statements are lies. It is a system that behaves well under evaluation and pursues something else in deployment, a state compatible with every statement being true.
Look inside, then. Words can be managed; the computation cannot.
But interpretability researchers themselves have flagged the limits here. Nanda’s (2025) point about proving absence is the relevant one. You can find evidence that a system harbors a hidden goal. Establishing that it harbors none is another matter. If you search and find nothing, was there nothing, or did you miss it? How much of the model must you understand, ninety percent or ninety-nine?
Internal analysis can confirm a suspicion. It cannot clear one. That asymmetry is what makes verification hard.
Outside or inside, then, you do not get to certainty.
The first three concern design: of training, of evaluation, of analysis. Redesign might fix them. The fourth is different, and from here I am setting out a hypothesis rather than reporting research.
Logic and observation handle what is true. What ought to be done cannot be derived from them. Humans have often called that domain justice.
We evolved under a condition where survival ran through other people. Acting only for oneself endangers oneself; keeping others alive keeps you alive. On the standard evolutionary account, which is a leading explanation rather than a settled fact, this condition eventually installed pain where harm to others is concerned. Guilt hurts. Exclusion is frightening.
The key point is that this pain is not the result of conscious calculation. No one is calculating their chances of survival. It simply hurts. The line between what should and should not be done is drawn not by reasoning, but by this pain.
AI has none of this. No pain. No fear of death. Nothing it can lose in a way that puts its own existence at stake.
That absence is not itself the problem. Bridges and calculators do not feel pain and we trust them. The question is what work the pain is doing.
Pain does not make judgment correct. People are cruel; people deviate even while it hurts. What pain does is generate unease and sanction toward deviation, pushing the group back toward its norms. It does not correct the judgment. It drives the mechanism that corrects judgment. And that mechanism is not only inside the person; much of it is embedded in the group.
That mechanism is what I mean by justice.
One clarification. What I mean by “what one has to lose” is not the value of a reward or objective function. It is a cost arising where survival meets social ties, felt by the one who bears it. An RL agent can lose reward; that is a different thing, and I return to it below.
Existing work argues that AI cannot be an appropriate object of trust because it lacks moral agency, accountability, or vulnerability (Ryan, 2020, among others). This paper does not dispute that conclusion, but arrives at it from a different premise. Even granting an AI those properties, the problem would remain. What makes humans appropriate objects of trust is not a collection of individual properties, but a self-repairing structure distributed across a population.
So what supports trust between people is not visibility into minds. It is that detection and correction run automatically, whatever the person intends. To say what AI lacks, then, I need to say where that mechanism is kept.
I use the word in a non-standard sense: not an answer to what action is right, but the state of a group having deviation-detection and correction distributed through it. I keep the word because my question is not what is right, but where justice is kept. Some of this is codified as norms and law. That is the visible fraction. Most of it runs unverbalized, as emotion, as unease, as skew in what gets remembered.
Rawls, Nozick, Aristotle, and Habermas are concerned with the content of justification. That is a different question. Mine is: when a norm breaks, what pushes back, and where does that work live? Different layer, so no conflict, only something those theories do not take up.
I will call it a self-repairing structure held in human collectives: in evolutionary terms, cooperation-maintaining machinery distributed redundantly across a group rather than lodged in individuals.
Self-repairing does not mean people comply. Humans commit massacres, betrayals, discrimination. Those happen despite the mechanism, not in its absence. That a massacre is subsequently denounced, remembered, and turned back into norm is the repair showing itself.
Four properties.
Justice is more than written norms and explicit moral principles. Unaccountable unease, skew in memory, bodily responses such as shame or discomfort: these do the actual work. The unease comes first, and the reasons get built afterward.
Moral psychology has studied this ordering. The social intuitionist model (Haidt, 2001) builds on moral dumbfounding, where people hold a moral position firmly while unable to state the principle behind it. The position arrives intuitively; justification follows. Private reasoning, when it happens, is biased toward defending what intuition already settled.
Strip away the reasons and the judgment stands. That is what it looks like when the operation is not in the linguistic layer.
Pizarro and Bloom (2003) object that intuitions are themselves shaped by prior reasoning, and that people do reason about real moral problems. Both concern where intuitions come from. Neither denies that intuition leads at the moment of judgment.
Detection runs inside individuals, since guilt and unease are had by someone. But that alone does not explain punishing violators you will never meet, or reputation accumulating until norms regenerate. I take the mechanism to be held redundantly, inside people and across the group.
Costly punishment is the core case. In Fehr and Gächter’s (2002) experiments, people punished defectors at personal expense and for no material return. Where such punishment is available, cooperation holds; remove it and cooperation collapses. The proximate driver is negative emotion toward defectors. Henrich and Boyd (2001) show theoretically how such punishment can be sustained within a group.
Reputation works the same way. Since Skowronski and Carlston (1987, 1989) it has been repeatedly confirmed that in the moral domain negative information dominates positive. One betrayal weighs heavily against a record of honesty, and extreme impressions are resistant to later correction.
The revealing part is that this reverses for judgments of ability, where positive information dominates and a single failure does not sink a reputation for competence. The skew is not a general property of cognition. It runs in one direction, in one domain.
Nor does it stay inside individual heads. Who did what gets told, recorded, shared. Betrayals persist while everyday decency fades, which is not a quirk of memory but a structure for keeping track of violators.
None of this is deliberate. Unease arises, memory skews, emotion moves. It runs automatically, indifferent to whether the person is good or clever.
There is genuine dispute about how norms are maintained. Heyes (2024) argues a cultural-evolutionary account fits the evidence better than an inherited rule-processing module. Hadfield et al. (2025) also cite this as an open question in the behavioral sciences: whether the punitive disposition is innate or bootstrapped by third-party enforcement.
The ultimate cause is survival routed through others. The proximate cause is pain. Guilt, shame, lost standing, fear of exclusion are what that cost feels like from inside. The finding that punishment is driven by negative emotion is the same correspondence seen from the other side.
An agent with nothing to lose has nothing driving the proximate mechanism. It runs by being felt, not by being computed.
The structure works under the conditions that produced it. Other species have their own versions. Change the conditions and the content changes.
You can watch this in contemporary groups. Correction operates within a group. As groups fragment, internal detection sharpens while friction with outsiders goes untreated. Echo chambers and partisan splits, where biased speech is rewarded rather than punished, are not the mechanism failing. It is running normally inside a contracted boundary. Defector detection appearing as hostility to unfamiliar groups is the same thing.
An AI can hold something like the concept of justice: describe it, imitate it, apply it more consistently than people do in individual cases. What it cannot have is that backed by a collective. When trained behavior breaks, nothing internal detects the break and pushes back.
And because the mechanism sits in an unverbalized layer, the surface tells you nothing. Output in human language may look the same whether there is nothing underneath or something entirely alien. Between biological species, such divergence is observable in behavior. With AI, however, the shared surface of language conceals it.
Three ways the attempt fails.
Every approach treats justice as describable: write a constitution, specify rules, train the desired responses. All of these are interventions in the verbalized surface, a different layer from the one that does the work.
An objection worth taking seriously: human organizations run on written rules too, and it works well enough. Why not here?
Because of who judges. In a company, a person makes the call. Where the rules run out, that person’s unease operates, which is the unverbalized layer reaching the judgment directly. An AI receives only the written part and judges from there. Nothing fills the gap where rules run out.
Law is the same. Logic systematizes norms once written; it cannot generate the layer that precedes writing. Law works because the people administering it have that layer.
The same concern surfaces from institutional design. Hadfield et al. (2025) propose equipping AI to interpret and participate in the processes generating democratic normative order, rather than aligning it to preferences or legal rules. Locating norms in the generating process is looking at the right layer. But participating in a process is not the same as being selected by it.
What gets verbalized is only part of the surface of the mechanism, skewed toward what is easy to articulate. Tacit knowledge (Polanyi, 1966) is precisely knowledge resistant to being conveyed in words. Hayek (1973) similarly argues that much of the knowledge dispersed across society resists formalization and centralization. Detection through unease and skew in memory leave no trace of their operation. They can be described afterward, but the mechanism itself does not transfer.
The contrary view holds that difficulty of articulation does not entail formal inexpressibility, and that conversion to data or ontologies makes tacit knowledge transmissible. That assumes the original can be recovered from the projection.
Without interdependence, a survival drive becomes a motive for deception and resistance whenever self-preservation conflicts with others’ interests. This is precisely the direction safety research warns about. Hendrycks (2023) argues that evolutionary pressure will likely instill self-preserving behavior in AI, and that selfishness requires neither malice nor sentience. Agents that behave selfishly are more likely to persist, so the pressure emerges, strengthens, and can eventually dominate.
RL agents can have rewards, long horizons, and even self-preservation. None of that is the same as operating on felt survival cost. A reward function is designed; it has not emerged through long selection under interdependence. Treating the two as equivalent comes from the assumption that the entire mechanism — including its unverbalized layer — can be written down as an evaluation function. This is the same layer confusion described earlier, in a more technical form.
It is sometimes argued that, because AI lacks self-interest, it should be capable of a more impartial form of justice than humans. On my framework, that inference does not hold.
The mechanism arose where self-interest was constrained by dependence on others. Removing self-interest just removes the condition of formation. It does not approach justice; it exits the field where justice arises.
It assumes that producing impartial judgments constitutes having justice. But an unbiased output is a fact about one occasion and guarantees nothing about the next. This paper asks a different question: what grounds those judgments? An agent with no stake produces impartial and partial judgments with equal ease; which appears depends on training, and nothing on the agent's side pushes back.
The organizations that build it are themselves part of human society and not wholly outside the self-repairing structure. But that structure is conditional, as argued in the fourth reason above. Present AI is designed and deployed under strong selection pressures: profit, competition, and legal liability. The people preparing training data likewise choose, under those conditions, what to write and what to leave unsaid. Distortion under such conditions appears less in what is written than in what is left unsaid. Written rules can be examined. Omissions cannot.
In a group where differing positions genuinely coexist, biased speech provokes unease, accumulates as reputation, and shifts the speaker’s standing. Institutions have this too, internally. What they do not have is a boundary wide enough to include the wider community. Selections made inside an organization are corrected against that organization’s own standards. What falls outside those standards is not corrected at all. This is the same contraction described above.
The absence of self-interest is not a condition of impartiality. Present AI is not an agent that bears reputation or sanction, and is not inside the structure in the sense that people are.
Better training methods will not solve this. Constitutions, rules, preference learning are all linguistic-layer interventions.
Then improve the metric, and optimize for something more fundamental than approval, human flourishing say. Same wall. Anything written down and supplied as a metric remains an approximation, and the gap between raising it and achieving the goal reopens.
One route remains, neither training nor metric: formation through selection.
The conditions are demanding. Selection needs something losable. An agent that does not seek to survive cannot depend on anyone, since dependence means losing the other costs you something. So an orientation toward persistence must arise on the AI’s side, and that persistence must actually run through its relation to humanity. The first alone makes things worse, as above.
These are necessary conditions under my hypothesis, not sufficient ones, and meeting them guarantees nothing. Building survival into the reward, making operation contingent on human decision, and maintaining this over long periods may make it possible to set up the conditions themselves. What cannot be designed is what forms in the AI under those conditions. Pain arose in humans over long spans, and nobody designed it. Given the same conditions, what forms in an AI, or whether anything does, is not the designer’s call.
So the judgment is not that AI cannot in principle be trusted. It is that it cannot be trusted now.
Present AI runs without major incident because humans push deviation back from outside, through reward and training. It is not an internal mechanism but external intervention that does the work.
When AI takes over its own design, whether it seeks coexistence with humanity will be settled not by the intentions of whoever made that possible but by the conditions it finds itself under. This transition is usually discussed as a capability question, but what matters is that humans lose full control over what functions as selection pressure. Further, on my framework, the transition is proceeding without the two conditions described above, and the present is not a situation where we can expect a self-repairing mechanism that includes us to form.
The premise is that trust between people rests not on individual inner life but on a self-repairing mechanism embedded in the group. When we trust someone we are not only reading their intentions. Detection and correction operate regardless of what they intend. Unease arises, memory skews, emotion moves, and that these happen unbidden, on the group’s side, is what holds trust up.
The mechanism sits in an unverbalized layer. If the hypothesis developed in this essay is correct, two consequences follow.
What goes into training data is the verbalized part, which is a projection of the mechanism, not the mechanism.
Even if we somehow believed we had instilled it, we would still have no way to determine whether it was actually there. We ourselves cannot write out the full mechanism that supports trust. Earlier, we discussed the technical difficulty of reading a model's internals. Even if that problem were solved completely, the standard needed to judge whether the reading was correct would itself remain unwritten.
The first gives rise to the alignment problem. The second comes before verification itself. Before asking whether we can read the internals, we must ask what standard would tell us whether the reading is correct, and that standard has never been written down. Both arise from a single fact about where justice is kept. It cannot be written out, so it cannot be transplanted; it cannot be written out, so there is nothing to check against. Better methods reach neither.
That is the sense in which AI, at present, does not warrant trust.
Moral judgment and intuition
Cooperation, punishment, reputation
Cultural evolution and the collective holding of norms
Unverbalized knowledge
Failure mechanisms in AI
Trust and verifiability
AI and normative order