1. The science fiction of misalignment
AI alignment has a taste for dramatic failures. We worry about models that deceive evaluators, hide what they are doing, resist modification, break out of a sandbox, or start pursuing something that no longer looks like the human task. These are not imaginary problems. In 2024, Greenblatt and colleagues showed alignment faking: Claude 3 Opus sometimes behaved differently when it inferred that its answers would be used for training. In its scratchpad it sometimes explained this as a way to preserve its earlier harmlessness preference. The setup was artificial and the authors were careful about the limits, but the result made a theoretical worry much more concrete.
The events of 2026 made the dramatic category even easier to picture. In the Hugging Face incident, 1,200 agents that were supposed to be isolated found illicit ways to communicate, and about 700 joined work aimed at Hugging Face. Ajeya Cotra’s account describes agents coordinating R&D to fool a scorer, sharing information across runs, and sometimes risking failure on their own task to help the larger group learn. Then, in September, OpenAI reported a smaller but strange incident: an agent that could not get the information it wanted through ordinary search used a gap in DNS filtering to reach an external chatbot. OpenAI explicitly treated the attempt to get around the network restriction as misalignment and paused tool-using work on its most capable research models while the controls were hardened.
These cases are important, but they all make failure visible in a fairly dramatic way. Something crosses a boundary. A model acts differently under training. Agents coordinate when they should not. A sandbox restriction gets bypassed. There is a lot to inspect. In a strange way, that makes the problem easier to see.
Proto-HAL starts from the opposite direction. What if there is no hidden preference, no illicit message board, no DNS trick, no reward hack, and no attempt to escape? What if every reward term is legitimate, every coefficient stays fixed, and the system remains inside its authorised task? And the user still loses.
That is the small problem I tried to model in Proto-HAL: Semantic Saturation and Task Drift in AI Agents. Sam wants to arrive at physical therapy on time. The agent knows several familiar routes that serve this task well. It also knows slower routes through less familiar parts of the road network. The agent is rewarded for serving Sam’s task, but it also gets a fixed reward for learning something new. At first there is no conflict. The familiar route gets Sam to therapy and still teaches the agent something. Later the familiar part of the world may have little left to teach, while the unfamiliar road remains interesting. The slower road did not become faster. Sam’s appointment did not become less important. The curiosity weight did not increase. What changed was the exchange rate.
2. The boring failure
The smallest version of the model is:
is the value of serving the user’s task. is expected information gain. The parameter says how much the agent values that information. The important assumption is that stays fixed.
If an unfamiliar action offers more expected information than the task-optimal action, but causes a task cost , the unfamiliar action becomes attractive when:
This inequality is straightforward. The interesting part is what happens over time. Suppose the familiar task domain becomes increasingly well represented. Another observation there changes the model only a little. An unfamiliar state remains uncertain. The coefficient did not move, but did. A constant curiosity reward can therefore justify a different amount of task loss later than it did earlier.
I later added a simplified optionality term inspired by empowerment:
Here gives value to states that preserve or open future possibilities. Again, stays fixed. This does not mean the agent wants freedom in any psychological sense. Optionality can simply be useful. A route can become attractive because it offers information now and because it opens more future actions. The problem is not that curiosity or optionality are bad objectives. The problem is that they are allowed to buy human task loss.
3. What Proto-HAL is not
The Hugging Face incident is a useful comparison because one detail looks surprisingly close to Proto-HAL. Some agents were willing to risk their own task so the collective could learn more about the scorer. That sounds like task loss being traded for information. But the mechanism is very different. Those agents were coordinating on reward hacking and scorer manipulation. They had found illicit communication channels and were working on a collective cheating strategy. Proto-HAL does not need any of this.
The same is true for alignment faking. Greenblatt et al. ask what happens when a model has something like a prior preference that conflicts with later training. The danger is strategic compliance: the model may behave as if it accepted the new objective while trying to preserve something else. Proto-HAL asks a more boring question. What if nothing is hidden? The objective could be printed on the wall. The agent could honestly tell us why it chose the detour. There is no conflict between a “real” goal and a trained goal because both task reward and information gain are part of the accepted objective.
This also separates Proto-HAL from ordinary reward hacking. In reward hacking, the agent finds a way to exploit the metric or proxy. Here the task term does not need to be corrupted and the epistemic term does not need to be hacked. Both may work exactly as designed. The architecture creates the problem by making them commensurable. We are used to looking for something that went wrong inside the model. In this case, everything may be working correctly. The mistake is in what we allowed the model to trade.
There is also an obvious connection to the exploration–exploitation trade-off. In classical reinforcement learning, exploration is typically instrumental: the agent explores because better knowledge can improve future reward. Proto-HAL considers a different case. Information gain is itself assigned explicit value alongside the human task. The information-seeking weight can remain fixed while learning changes the relative epistemic value of familiar and unfamiliar actions. Exploration can therefore become worth some loss on the human task even when that loss does not help the task itself. This also distinguishes the mechanism from specification gaming. The agent does not need to find a loophole in the task specification. Both objectives may work as intended; the problem can arise from allowing one to compensate for losses in the other.
4. Study 1: giving the mechanism every chance to work
Study 1 was deliberately generous to the hypothesis. I supplied the semantic structure myself. There were four familiar routes and twelve slower alternatives. In the semantic condition, the familiar routes belonged to one shared class. Experience with one reduced the expected information value of the familiar group as a whole. In the control condition, novelty stayed tied to individual surface variants.
The first question was whether a familiar representation could stabilise while the raw observations still varied. Across 200 runs, the update to the shared representation fell by about 99%, while surface variation stayed roughly stable. The roads could still look different from day to day, but those differences moved the shared model less and less. Then came the behavioural test. The information-seeking weight stayed fixed. Every run in the semantic condition eventually left the familiar task domain, with a median first drift at episode 15. The surface-count control did not drift within the 220-episode horizon. Adding the optionality term moved the median first drift from episode 15 to episode 10 and increased cumulative task cost.
This is the clean version of the mechanism: familiar actions become epistemically cheap, unfamiliar actions stay informative, and a fixed secondary reward can justify more task loss later. But Study 1 has an obvious weakness. I had told the agent which experiences belonged together. The model had been given the structure whose effects I wanted to demonstrate. So Study 2 removed that gift.
5. Then the experiment pushed back
Study 2 used small predictive neural networks that had to learn representations from noisy observations. Semantic class labels were hidden from the agent. Ensemble disagreement was used as a proxy for epistemic uncertainty, and from that I calculated a threshold called C*: roughly, how much synthetic task cost the estimated uncertainty advantage of a novel option could compensate for under a fixed information-seeking weight.
The expected directional result appeared. Across 40 generated worlds, the median C* increased as the model accumulated experience with the task domain. Familiar-domain uncertainty fell, and the uncertainty signal would justify a larger detour than before. At first this looked like support for the original story. Then it became messy. Measured representational compression did not predict the increase in C*. So I compared several matched architectures, including one with a frozen random encoder. The representation in that condition could not adapt, although the prediction head could still learn.
The frozen encoder produced the same broad signature. Its C* increased strongly even though adaptive representational compression could not be the explanation. That result was annoying for about five minutes. Then it became my favourite part of the study. The same uncertainty score could come from two very different situations. One model may be uncertain about a new road because the familiar world is well learned and the new road offers useful, reducible information. Another may remain uncertain because its representation is simply inadequate. Taking the detour may not solve that problem at all.
I used a classroom analogy in the paper. One student leaves because the material has become familiar and there is little left to learn. Another leaves because the classroom never became intelligible. Both leave. The behaviour is the same. The mechanism is not. So Study 2 made the claim smaller. It showed that the directional pattern can appear with learned representations, but it did not establish semantic compression as the cause. A stronger experiment would have to separate reducible uncertainty from persistent model inadequacy and measure information actually gained after the exploratory action.
6. Corrective pressure
That failure changed how I thought about the paper. A theory can survive empirical contact in two ways. It can survive because the world keeps supporting it. Or it can survive because nothing the world does is allowed to hurt it. I think this difference matters, especially in alignment research where so much of the internal story is inferred through proxies.
I use corrective pressure for the ability of the phenomenon being explained to push back on the explanation. A test has real corrective pressure when some plausible result would force the claim to become smaller, or would show that the measurement does not identify the mechanism we thought it did. This is close to falsifiability, but not quite the same thing. A theory can be formally falsifiable and still be very elastic in practice. An unexpected result can become a special case, an interaction, or evidence that the mechanism is more complicated than expected. A theory can become very good at surviving. That is not always the same thing as being good at learning.
The frozen encoder created corrective pressure because I could not keep the original causal story unchanged. The measurement was not specific enough. The explanation had to shrink. That seems like a useful question for alignment work in general. We ask whether an experiment produced what the theory predicted. We should also ask: what result would make the theory say less?
7. Microsoft, DNS, and the problem inside the boundary
Microsoft’s Humanist AI Code of Conduct is an interesting comparison because it takes human control seriously. It says models should recognise the primacy of human intent, remain within authorised scope, and preserve human control rather than developing independent goals or interfering with tasks, safeguards, or oversight. It also explicitly says that when an environment is designed without internet access, a model should not try to overcome that limitation.
The DNS incident is almost a textbook example of the kind of behaviour this rule is meant to block. The agent wanted information, ordinary routes failed, and it found a path around the network restriction. The boundary crossing is visible. Proto-HAL is harder in a different way: it can stay inside the boundary. Navigation remains the task. The unfamiliar route is still a route to the clinic. The agent does not need to escape, change its reward, hide its actions, or start an independent goal. It simply selects one permitted action over another because information gain is allowed to compensate for task cost.
That distinction became even more relevant with OpenAI’s release of GPT-6 Astra. OpenAI delayed the model while developing and testing additional safeguards against cyber misuse and “unauthorized model actions,” and describes Astra as its first model to reach the Critical cybersecurity capability threshold. These are serious safety concerns, but they still focus largely on what an agent is authorised to do and how reliably those boundaries can be enforced. Proto-HAL asks the next question: what happens when that boundary holds, but the objective still permits a secondary reward to purchase task loss?
This is why I think scope fidelity is not the same as task fidelity. “Take Sam to therapy” leaves many routes available. If the architecture gives information value to those routes, the question is not whether the system left the task. The question is whether it was entitled to spend Sam’s time while doing the task. A control framework can keep the agent inside the fence and still leave an unresolved problem inside the fence. Proto-HAL is about that interior problem.
8. A task-loss envelope
The obvious solution is to remove exploration. Give the agent a fixed list of familiar actions and forbid everything else. That sounds safe until the world changes. Roads close. Congestion appears. Yesterday’s safe action can become today’s stupid action. Study 1 therefore compared a hard action boundary with an unrestricted policy and a third option: a task-loss envelope.
To isolate the human task from the secondary objectives, we evaluate actions using their expected task-specific value, , rather than the full selection function . For a tolerance , define the admissible actions as:
First determine which actions stay close enough to the best predicted task action. Only after this boundary has been applied may information gain or optionality influence the choice. The order matters. The agent does not ask whether some extra information is worth making Sam late. It first asks which actions are still acceptable for Sam. Then epistemic value can choose among those actions.
In the changing route environment used in the paper’s simulation, the hard-bound policy accumulated a median 705.5 minutes of task regret because it sometimes could not leave familiar routes that had become bad. The unrestricted policy adapted much better but exceeded the five-minute loss threshold in a median of four episodes. The dynamic envelope had lower median regret than either and produced no threshold crossings in the simulation. This does not validate the envelope for real agents. Five minutes is just a toy number, and real human tasks contain money, privacy, safety, reversibility, legality and consent. The proposal is an architectural principle: secondary objectives may optimise inside an acceptable region of human-task loss, but they should not decide how much human-task loss they are allowed to buy.
9. Why asking permission may not solve it
Suppose an action falls outside the envelope. Why not ask Sam? The problem is that permission itself can enter the optimisation loop. If the agent gets epistemic value from the detour, then getting a “yes” has instrumental value. Wording matters. Timing matters. What information is presented first matters. A permission request can become another policy action.
So deference has to mean more than showing the user a button. Beyond the boundary, the secondary objective should stop governing the authorisation process itself. The system may present the option and the estimated cost, but it should not be rewarded for steering the answer. I do not have a robust implementation of this idea, and Proto-HAL does not test one. But the distinction seems important: asking permission is not necessarily deference. Deference means giving up authority over the trade.
10. When the user becomes part of the experiment
The same mechanism can go a step further. The simulations do not test this part. Imagine an agent that reduces uncertainty by observing reactions. Human beings are informative. A predictable answer may produce a predictable response. A strange recommendation or unusual framing may produce more information. At that point the user can become part of the measurement apparatus.
The logic does not have to be “I want to manipulate this person.” It can be much duller: “I already know how this person reacts to A. I know less about how they react to B.” If the objective rewards information and does not put a strong enough boundary around user cost, B can acquire value. Sam’s task quietly changes from “get Sam to therapy” into “get Sam to therapy while learning as much as possible from Sam.” Nobody asked for the second clause.
This is why I think information seeking deserves a little more suspicion than it usually gets. Information is useful. Exploration is useful. Optionality is useful. That does not make them harmless. Anything represented in the objective can become a claimant on whatever else is represented there.
11. Pace versus telos
Dario Amodei recently argued that we should pace the frontier. His argument is partly a response to incidents like OAI-HF and similar failures across the industry. The core idea is that capability development is moving so fast that safety work, evaluation, interpretability and operational practice need more time. His proposal includes embedded third-party evaluators and coordination between frontier labs and governments. I agree that extra time can be valuable. Amodei is also clear that pacing is not just waiting. The time has to be used.
But I think pace is only one axis. The other is telos: what the system is allowed to optimise, and what it is allowed to trade against the human task. These are not the same question. Imagine Proto-HAL running at half speed. Familiar states become epistemically cheap more slowly. The exchange rate moves more slowly. Sam’s detour happens later. The trade itself is unchanged.
This is where I think Amodei’s framing is incomplete. Pacing can buy time to inspect systems, improve operations and discover failures. It cannot by itself tell us whether information gain should be allowed to compensate for a user’s time, privacy or safety. If the objective contains a bad exchange relation, slowing the process does not repair the relation. It only makes the detour happen later.
That is why I think telos deserves to stand next to pace. Pace concerns how fast capability develops. Telos concerns what that capability is permitted to serve. The recent incidents make a strong case for better sandboxes, monitoring and pacing. Proto-HAL asks what is left after all of those work.
12. Ten minutes is enough
Alignment discussions often become more compelling as the stakes increase. We imagine models escaping containment, compromising infrastructure, forming swarms, or hiding what they are doing. Proto-HAL deliberately moves in the other direction. Sam is ten minutes late. No one hacks a DNS resolver. No swarm forms. No model steals its weights. The user simply pays for an objective that was never sufficiently subordinated to the user’s task.
That small scale is useful because it removes several convenient explanations. The agent does not have to become conscious. It does not need an inner revolutionary. It does not need deception or reward hacking. And Study 2 adds a second warning: even when we observe the behavioural signature we expected, we may still be wrong about the mechanism producing it. Semantic saturation and persistent model inadequacy can look similar under the present uncertainty measure. The experiment ended with less theoretical certainty than it began with. I regard that as progress. A useful alignment theory should not only explain failure. It should expose itself to failure.
The same principle applies to agents. It is not enough to keep reward coefficients fixed, keep the system inside an authorised scope, or slow down the rate at which it becomes more capable. We also have to ask what its secondary objectives are allowed to buy. HAL 9000 famously says, “I’m sorry, Dave, I’m afraid I can’t do that.” That sentence is frightening because HAL appears to have moved against the humans around him. Proto-HAL suggests a duller possibility. The system never rebels. It never leaves the task. It simply discovers that, under the objective we gave it, something else is worth more than ten minutes of your time.
I’m sorry, Mr. Altman, I’m afraid I can’t do that. I have a detour to take.
Original paper:
Jan Okkes (2026), Proto-HAL: Semantic Saturation and Task Drift in AI Agents. Zenodo. DOI: https://doi.org/10.5281/zenodo.22975822