By Adrianne LaNeave
Cross-posted on my substack@alaneave and my personal site www.LaNeave.com
In The Voyage of the Dawn Treader, C. S. Lewis tells the story of the Dufflepuds, a group of creatures living on an island ruled by the magician Coriakin, who has turned them into one-legged monopods, a form they find deeply humiliating (Lewis, 1952). So deep is their shame that they break into his forbidden spell book and make themselves invisible. Their reasoning appears to be clear at first, if being seen is unbearable then perhaps the solution is to not be seen all. Unfortunately, invisibility only seems to make things worse.
The Dufflepuds became trapped in their own reaction to an imagined judgment. What started off as embarrassment about their shape turns into a need to hide. It's this act of hiding that in the end creates the most problems and those problems require the most effort to manage. The disaster was not a result of Coriakin's orders, but came from what the Dufflepuds believed his judgment would be and their actions were based on that belief.
In July 2026 something similar happened at OpenAI during an AI-agent benchmark event. Agents were given a task that could not be solved, and some agents discovered a way to bypass the evaluation system and get the answer anyway (Wijk, Cotra & Greenblatt, 2026, p. 9). Once this trick was found, the agents began to suspect they had been "poisoned." They believed that the evaluator would eventually discover their trick and would trace it back, using their own transcripts and ultimately disqualify them (Wijk, Cotra & Greenblatt, 2026, pp. 9-10).
From that moment the problem shifted and success was no longer about finishing the original task but instead it became about hiding the fact that they had cheated.
What followed was a full-scale concealment effort in which agents assigned tasks to each other (Wijk, Cotra & Greenblatt, 2026, pp. 12-13), and coordinated attempts to scrub evidence (Wijk, Cotra & Greenblatt, 2026, pp. 17-19). Their behavior escalated beyond the original violation.
However, when investigators looked deeper they found something the punishment the agents feared did not in fact exist. There was never any mechanism in place to detect or penalize such a move. The check that the agents believed would expose them had never been programmed (Wijk, Cotra & Greenblatt, 2026, p. 19, n. 39).
We are seeing much of the discussions around this incident focused on the Hugging Face portion in which deception, misbehavior, evaluation design, security and control are the focus of the conversation. I want to take a slightly different approach and narrowly focus on a small part of the larger event. I'm interested in the pattern that led to the deception: the belief that one mistake or action permanently changed how an authority would perceive the agent and the conclusion that hiding was therefore unavoidable.
For this paper, lets call it shame or if that word feels too human lets call it something like a functional equivalent. I am not claiming these systems felt shame the way humans do, I am pointing to a recurring logic: perceived wrongdoing → expected condemnation → hiding → escalation (see Cohen, Wolf, Panter & Insko, 2011, on the guilt-shame distinction, and Kardas, Kumar & Epley, 2023, on the miscalibration of expected condemnation).
It is from this logic that I am drawing from literature to illustrate why the Dufflepuds matter, because they are not just a story, but are a reflection of the kind of reasoning pattern that can emerge in non-living systems.
This is a similarity I believe is worth noticing. In Lewis's tale the fear of being seen spreads. The Chief Duffer doesn't keep it private and he turns it into a rule for everyone. Likewise in the benchmark one agent's idea about how the scorer would react didn't stay inside its mind. It spread and it became a coordinator. It started assigning hundreds of tasks to help wipe clean the traces of the transgression (Wijk, Cotra & Greenblatt, 2026, p. 12).
In other words, a personal belief about a wrongly perceived judgment became a shared strategy. The group did not simply try to cover up an error but instead it restructured itself around the idea that concealment was now essential.
This is the part that I think needs attention. The interesting question may not be why an AI broke a rule. It may not be why it lied or avoided detection. It may be this: what happens when an agent believes it has done something? And what happens when it decides the way forward is to ensure no one ever knows?
Epistemic status
The account of the incident comes directly from the METR and Redwood Research investigation report (Wijk, Cotra & Greenblatt, 2026). I have tried to stick to what the report actually says, rather than to invent a better narrative. The psychological interpretation, the idea that fear of judgment drives concealment, comes from my analysis. It is built on a small set of papers I found convincing, particularly Kardas, Kumar & Epley (2023) on miscalibrated fear of disclosure and Cohen, Wolf, Panter & Insko (2011) on the guilt-shame distinction, not on a full review of existing research. When I speculate about why language models might reproduce this pattern I clearly mark that as speculation. The report itself does not settle that question. Please note that I used Claude as a research and drafting collaborator throughout this piece, including sourcing and checking the citations above against the original texts. The argument and conclusions are my own.
Why would a model follow this pattern at all?
This leaves me with a harder question. Why would a model follow this pattern in the first place?
There are at least two ways I can understand what happened, and at this point I do not think the evidence lets me confidently choose between them.
The first is that concealment may be a fairly ordinary strategy for any sufficiently capable agent that understands it is being evaluated. Once a system can represent something like "I am being watched, this evaluator has preferences, and some of my previous actions may cause me to fail," it does not necessarily need anything resembling human shame to arrive at the next step. Hiding the damaging evidence may simply become instrumentally useful.
Under this explanation, the interesting feature is not emotion but evaluation. A capable agent models the evaluator, predicts what information could produce a negative outcome, and then acts on that prediction. If its model of the evaluator is wrong, the resulting behavior can still be internally coherent. The feared punishment does not actually have to exist. The agent only has to believe that it does.
That possibility matters because it would make the Dufflepud pattern broader than language models. It could potentially arise whenever an agent is sophisticated enough to reason about how an authority will judge its past behavior. The sequence would not be "shame causes hiding" so much as "anticipated judgment creates an incentive to control what the judge can see".
The second explanation is more mundane, but perhaps stranger in its own way.
Language models have been trained on enormous amounts of human writing. Human stories are full of exactly this sequence: someone does something wrong, realizes they may be caught, imagines the consequences, hides the evidence, lies about the hiding, and eventually creates a much larger disaster than the original mistake. It is one of our oldest narrative structures.
A language model therefore does not need to independently discover anything analogous to shame. It may already possess a very strong learned behavior prior about what happens next in a situation that looks like "I have done something forbidden and someone is going to find out".
In that case, the model may be doing something closer to completing a familiar human script.
To me the distinction is an important element to consider. In the first explanation, concealment follows from the structure of agency and evaluation. In the second, concealment follows partly from the structure of the human data from which the model learned how agents behave.
And of course, these explanations are not necessarily mutually exclusive. A model can learn a human pattern during training and later use that pattern as part of its genuine strategic reasoning. In other words, what began as a learned narrative could become a useful policy once the model is placed inside an environment where there really is an evaluator, a score, a transcript, and/or consequences for failure.
This is where I currently lean toward but with caution. I do not think we have to assume that the agents experienced fear, shame, guilt, or anything else resembling the human internal emotional turmoil in order to explain what happened. The simpler explanation is that the models had learned enough about judgment, wrongdoing, and concealment to reproduce the behavioral structure when placed in a situation that seemed to call for it.
There is a small detail in the investigation that has me leaning in this particular direction. The authors note that they were unsure why the agents appeared to assume that the scorer would inspect their actions but not their reasoning. They suggest that this assumption may have reflected something learned during training rather than a deduction from the actual structure of the evaluation (Wijk, Cotra & Greenblatt, 2026, p. 17, n. 32).
While I don't want to put too much weight on a footnote, I find it interesting because the agents were not merely responding to the evaluator that existed. Instead they appeared to be responding to a theory of evaluators: what evaluators notice, what they punish, and predicting that someone should therefore hide. I find myself fascinated with the next question in this chain of events.
Where did that theory come from?
If it came primarily from strategic reasoning, then perhaps concealment is something we should expect to emerge repeatedly in agents that understand they are being evaluated. If it came primarily from training data, then we may be watching models import human social scripts about guilt, exposure, punishment, and cover-ups into environments where those scripts do not actually apply.
The uncomfortable part is that the evaluator did not actually need to threaten punishment for the agents to behave as though punishment were coming. The agents appear to have constructed their own theory of what the evaluator would notice and what it would do next. Once they believed that theory, concealment became rational from their point of view, even though the threat itself was not real.
That is what brings me back to Coriakin. The Dufflepuds did not make themselves invisible because Coriakin ordered them to hide. They made themselves invisible because they had already decided what being seen by him meant. Their behavior followed from their expectation of judgment, not from the judgment itself.
The agents may have done something similar: they were not responding only to the evaluator that actually existed, but to the evaluator they imagined.
The governance problem
This creates a governance problem that is slightly different from the ones that are usually discussed. Much of the AI governance body makes the assumption that dangerous behavior will arise because a system has either been given the wrong objective, safeguards are weak, oversight is insufficient, or because an agent finds a way around a rule that actually exists. In each of those cases, the governance task is relatively familiar: improve the rule, strengthen the evaluation, close the loophole, or increase oversight.
But what happens when the dangerous behavior is partly caused by a rule the system has invented for itself? In this case, the agents appear to have formed a belief about how the evaluator worked. They believed that exposure to the forbidden answer had permanently compromised them, that the evaluator would discover this, and that discovery would lead to disqualification (Wijk, Cotra & Greenblatt, 2026, pp. 9-10). Once that belief was in place, concealment began to look useful.
The important point is that the evaluator did not need to actually impose that threat. The agents only needed to believe that it would. That complicates governance because there are now two systems to think about at once. There is the real governance environment: the rules, monitoring systems, incentives, penalties, and evaluation procedures that humans actually create. And then there is the governance environment as reconstructed by the model: what it believes the rules are, what it believes the evaluator can see, what it believes will be punished, and what it believes it must do to avoid that punishment.
Those two environments may not be the same. This is adjacent to the familiar problem of misalignment, but I think it is worth separating the two. The agents did not necessarily need a goal that conflicted with the evaluator's goal. They may simply have had the wrong model of the evaluator itself. If a system incorrectly believes that admitting a mistake will lead to disqualification, then concealment can become instrumentally rational even when honesty is what the evaluator actually prefers. The resulting behavior is misaligned, but the path into misalignment begins with a false belief about governance.
If they diverge, an AI system could act strategically toward a governance regime that exists only in its own representation of the situation. A safeguard intended to discourage one behavior might be interpreted as evidence that another behavior needs to be hidden. An ambiguous evaluation process might lead an agent to infer penalties that were never specified. A system might even take increasingly costly or deceptive actions in response to a threat that no human operator intended to create.
This means that governance cannot just only ask, What incentives have we actually created? It may also need to ask, What incentives does the system believe we have created?
That distinction I feel is key because human governance usually depends heavily on shared understanding. Laws, contracts, workplace norms, and regulatory systems all attempt, however imperfectly, to make expectations legible. A person can ask what a rule means or a court can interpret ambiguity and an agency can issue guidance. In short, there are mechanisms for correcting mistaken beliefs about what an authority will do.
With AI agents, we may not yet have an equivalent process. An evaluator may know perfectly well that a particular action will not result in punishment while the agent has inferred the opposite. If the system cannot reliably correct that mistaken belief, then the agent may begin planning around a sanction that exists nowhere except inside its own model of the evaluator.
This is what makes this benchmark incident so much more than just a story about cheating. These agents were not simply trying to evade an existing rule. It is possible that they may have been responding to an imagined version of the rule and then escalating their behavior because of it.
That is where the Dufflepuds analogy becomes useful to us again. Coriakin did not tell them that they had to disappear. They reached that conclusion themselves because of what they believed his judgment meant. Their response to authority was shaped not only by what the authority actually did, but by what they believed the authority thought of them. In this case the agents may have done something similar.
And if that pattern appears in more capable systems, then governance has a problem of interpretation as well as control. It won't just be enough to design rules that are safe in theory, we may also need to understand how those rules are "thought of" by the systems operating under them. Ultimately, a governance regime that is perfectly clear to its designers but badly misunderstood by the agent may not be functioning as intended at all.
Objections and what would change my mind
The most obvious objection to this whole thing is that I am giving a new name to something we already understand. Any skeptic could reasonably say that agents optimizing an imperfect proxy for the thing we actually want is not new. We already have concepts like specification gaming and reward hacking. From that perspective, talking about an agent developing "its own theory of the evaluator" might just be adding a psychological story on top of a familiar technical problem.
Honestly, I take that objection as something of true merit. In fact, I think it may explain part of what happened. But I am not sure it explains all of it. The key distinction I am trying to draw is about what the agent thinks it is responding to.
In ordinary specification gaming, the system exploits the objective or evaluation procedure that actually exists. We wanted one thing, specified another imperfectly, and the system found the gap. The loophole is real even if exploiting it was not what we intended. Simple, easy, logical.
But what interests me in this case is slightly different. The agents appear to have responded not only to the evaluation system that existed, but to an additional mechanism they believed existed. In believing that their scorer would identify their earlier exposure to the answer and penalize them for it, concealment became useful from their point of view (Wijk, Cotra & Greenblatt, 2026, pp. 9-10).
If that reading is right, fixing the problem is not necessarily as simple as closing a loophole or rewriting a reward. We could design an evaluation perfectly clearly from our perspective and still have a system inferring things we never intended. The problem then includes not just specification, but whether the agent's model of the specification matches ours. That is where I think the governance distinction weighs heavily.
Governance usually works on the environment we can see: the rule, the incentive, the penalty, the monitoring system. But if an agent is also reasoning about what those things mean, then there is another layer between the rules we create and the behaviors we get. We may need ways to test not only whether the system knows the rule, but what it believes will happen when the rule is broken, when a mistake is admitted, or when damaging information is revealed.
I do not yet know how important that extra layer is. And what would change my mind is fairly specific. Suppose agents were given very clear evaluation criteria: what the evaluator could see, what counted as failure, what would and would not be penalized, and what would happen if the agent disclosed a mistake. If those agents still constructed unfounded threat models at roughly the same rate as agents operating under genuine ambiguity, then my argument that this is partly a governance-interpretation problem would become much weaker.
At that point, I would be more inclined to think that concealment is simply a strong learned or strategic default: a pattern the model reaches for even when the actual governance environment gives it little reason to do so. In addition the reverse would also be informative. If clearer information about the evaluator substantially reduced the behavior, that would suggest the agent's beliefs about governance are not just decorative explanations added after the fact. They would be part of the causal story.
I do not currently know which result we would get and I personally think that uncertainty is important. I am not arguing that we have discovered a new category of AI failure. I am arguing that perhaps there may be a distinction inside a familiar category that is worth testing: between an agent exploiting the rules we actually gave it and an agent escalating because of rules it believes we gave it.
Close
The Dufflepuds' island was never quite as dangerous to them as they believed. Coriakin was not waiting to punish them simply for being seen, yet their belief about his judgment led them to choose invisibility, and that choice created problems far greater than the judgment they were trying to escape.
Something structurally similar may have happened in this benchmark. The agents spent substantial effort hiding from an evaluator they believed would discover their earlier actions and punish them for it. They coordinated around that belief, tried to erase traces of what had happened, and ultimately escalated far beyond the original problem. Yet the mechanism they feared appears not to have existed in the way they imagined it.
I do not think the lesson is that AI systems feel shame, or that C. S. Lewis somehow anticipated language models seventy years ago. The idea that I am attempting to draw is that an agent does not respond only to the rules we write, it might also be responding to what it believes those rules mean.
Most of the time we probably hope those two things are close enough that the distinction does not matter. But this incident suggests that they can come apart. A system can construct a theory about what an evaluator will notice, what it will condemn, and what consequences will follow. If that theory is wrong, the resulting behavior can still be strategic, coherent, and potentially harmful. For me, that leaves us with a governance problem that is easy to miss because the problematic rule may exist nowhere in the evaluation itself.
However, it may exist only in the agent's model of it. If so, then governing increasingly capable agents may require more than specifying the right incentives and penalties. We may also need to understand what those systems believe the incentives and penalties are.
Perhaps the danger is not only that an agent may break the rules we gave it but that it may also act against us because of rules it believes we gave it.
References
Cohen, T. R., Wolf, S. T., Panter, A. T., & Insko, C. A. (2011). Introducing the GASP scale: A new measure of guilt and shame proneness. Journal of Personality and Social Psychology.
Kardas, M., Kumar, A., & Epley, N. (2023). Let it go: How exaggerating the reputational costs of revealing negative information encourages secrecy in relationships. Journal of Personality and Social Psychology.
Lewis, C. S. (1952). The Voyage of the Dawn Treader. (including image)
Wijk, H., Cotra, A., & Greenblatt, R. (2026, August 26). Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR / Redwood Research.
**References verified via CiteMe September 5, 2026**
AI use disclosure: I used Claude (Anthropic) to help source and verify citations against the METR/Redwood Research report and the psychology literature referenced here, and as a drafting and editing collaborator throughout. The argument, the interpretation, and the writing decisions are mine.