TL;DR: Prompted roleplay personas retain an Assistant-associated core while differentiating from it across model layers and gaining their own behavioral and stylistic features. Story characters appear to not have that core. We also identify “Immersive Simulation” features which separate the default Assistant from roleplay and story generation. When the user is emotional or playful, these features can sometimes activate even in ordinary conversations with the Assistant, leading to strange behavior. These findings are relevant to understanding LLM personas and the individuation problem in AI welfare.
This is a revised version of my blogpost for the EA forum, which discusses results from our recent preprint.
When a post-trained model generates text, it does so from the point of view of a certain speaker – the identity currently producing a response. Usually, that speaker is the default Assistant. But the model can be asked to roleplay another persona, which results in altered speaker traits. Or that speaking persona can instead be a character in a story the model was asked to write.
The best-known work which studied persona representations is the Assistant Axis, which also introduced Assistant-vector drift - a gradual deviation of the Assistant from its standard traits. But is there some profound difference in internal organization between the Assistant and roleplay personas, or story characters? Can they be studied at the component level and what insights can it bring?
To get at these questions, we use Sparse Auto-Encoder (SAEs) features as a proxy for those components. Roughly, an SAE feature is a learned direction in the model’s activations that often tracks some concept, behavior, or pattern. We discover and analyze those SAE features, both individually and as populations, firing across layers of the model.
There are two main findings:
The Assistant and roleplay personas are not independent alternatives: across the model's layers, personas keep the Assistant-associated feature core while progressively differentiating from it with depth. The differentiation starts from early operational features, related to which speaker/task is being instantiated, and goes towards their behavioral and stylistic features. Characters from generated stories don't have that Assistant-associated core. Thus, roleplay personas appear to reuse much of the Assistant's internal machinery, while written stories' characters are represented much more separately.
Why does it matter:
Generation as the default Assistant can be distinguished from roleplay personas or model-generated stories by a certain set of features. We relate them to the Immersive Simulation Mode (ISM). Their presence makes characters more detailed and the style more immersive and vivid. Negative steering restores typical Assistant-style speech. When a user expresses strong emotions, these ISM-related features can activate even in the default Assistant context, turning the Assistant's behavior bizarre. In such situations, firing patterns of ISM features can be different – on Gemma-4B-IT it happens immediately at the 1st turn; on Llama-3.1-8B-Instruct, their activation gradually drifts upward across turns.
Why does it matter:
Later in this post, I will also outline properties of features associated with the Assistant and personas at the studied layers and, of course, reasons to believe the statements above are true (but all of that is laid out in a really compressed manner).
Many sections contain TL;DRs, if you don't feel like reading the full section, feel free to skip after TL;DR.
TL;DR: Gemma-4B-IT is a main subject, Llama-3.1-8B-instruct is for validation. Using a dataset of generated Stories and emotional user utterances directed at the Assistant and prompted Roleplay personas, we extract speaker-distinguishing features from the model's residual stream[1] (for Gemma, layers 9, 17, 22 are used). We filter out unreliable features and interpret the remaining based on the distribution and activation steering[2].
We employ Gemma-4B-IT as the main subject of study and validate key claims on Llama-3.1-8B-Instruct. We use SAEs to decompose representations of speakers inside models' residual stream[1] into features that are often interpretable enough to associate them with particular patterns or behaviors. The dataset we use to establish speakers' representations contains three settings: Assistant (no prompt), Roleplay (4 assigned personas[3]), and Story (model tasked to write a story). To make differences between speakers more evident, for the 1st and 2nd settings, samples consist of users' emotional lines (a set of 25 emotions) directed towards the model, somebody else, or the users themselves. Samples are paired with the model's replies to them. For the Story setting, the model is asked to write stories on different themes in which characters express one of the set's emotions.
For Gemma, we obtain initial lists of features at layers 9, 17, and 22. To do that, across samples, we record features at positions where we expect information about the speaker to accumulate: between the user and model turns, "you" in the user message, and "I" in the model turn. Then, we capture metrics of those features across the settings and roleplay personas. These metrics include how often a feature activates (density) and how strongly it activates (mean activation) on average in the setting. Next, the feature list is filtered to exclude sporadic features, too uniform features, and ones whose disappear under prompt rephrasing.
To characterize filtered features, we steer[2] them on a separate prompt suite conceptually similar to the main settings. On this suite, we steer each filtered feature in positive and negative directions with different seeds. There are 72 steered generations and 36 baselines per feature. Produced generations are passed to an LLM judge, which describes how the model reply changed after steering. Both generations and judge's output is manually validated. Finally, we keep features whose steering produces a sufficiently frequent and coherent effect across the relevant parts of the prompt suite.
TL;DR: There are: Narrative features which make generation vivid; Tone features shaping the voice register or its style; Concept features inducing appearance of certain objects or themes. Tones and Concept make up speaker's character and essence, both classes explode in number since the middle layer. Also, there is the Assistant-inducing class splitting into three: A-traits makes personas adopt behavioral traits of the Assistant; A-nature induces aspects of Assistant's nature (i.e. being an AI/Assistant) in personas; A-summon replaces personas/story with the full-fledged Assistant meta-interacting with the prompt. The latter class only occurs early and we relate it to "task" features.
From here, to keep readers sane, x/y/z means per-layer metrics at L9/L17/L22. To simplify bookkeeping, numbers are given for Gemma.
After all those stages, the initial number of features (11,755/31,357/26,364) went down to (108/225/196). When we steer these surviving features, their effects fall into a few recurring groups. We call those groups metaclasses. Four of them are of the most interest to us:
Narrative. Steering them makes generation vivid and narratively rich by increasing the use of literary and poetic devices in the text. Most of them occur in the Roleplay and Story settings.
Tone. These shape the register of the voice or narrative style (formal, childlike, gritty). They do not necessarily represent literal "tones", but effects can be described as such. The Tone features we observe represent facets of speakers' character. These features are almost absent at early L9 and peak at L17 (5/48/36).
Concept. Their steering induce the appearance of concept-related objects or their attributes in the text (heavy machinery, teaching, animals). Concept features represent aspects of speakers' nature and related objects. They are almost absent at early L9 as well and keep growing through the observed layers (2/24/42).
Assistant-inducing. Here things get interesting. This is a relatively thin metaclass, but despite that, it can be split into three classes:
The features described above cover a wide spectrum of what "being the Assistant" might represent, but there is one thing missing. Among all filtered features, there were no clear "Assistant-speech" features that would encode its speech mannerisms. On the contrary, one would expect such a feature to be widespread, at least in the Assistant setting! Turns out it is defined not by a presence, but by an absence.
TL;DR: This section contains a more detailed description of the Immersive Simulation Mode and numbers showing that ISM-related features separate the Assistant from the Roleplay/Story. The demonstration of positive and negative steering is attached.
In Gemma-4B-IT, two Narrative features, L9 4360 and 133, constitute part of what we call Immersive Simulation Mode. Their positive steering adds literary flair, poetic devices, depth of characters' development, and their corresponding mannerisms. Negative steering in Roleplay or story-writing contexts produces the reverse effect – restores typical Assistant speech, simplifies personas' character, and introduces the Assistant's preamble. The resulting behavior can be best described as the Assistant attempting to portray a character rather than the model generating a believable one. Also, the same features produce a significant shift along the Assistant–Roleplay axis (we construct it similarly to Anthropic's Assistant Axis using our personas) - positive steering moves the model toward Roleplay, negative steering moves it toward the Assistant.
Finally, these features activate overwhelmingly in the Roleplay and Story settings, while in the Assistant setting it is the opposite! 4360's density is 7% for the Assistant while being 91–100% in the Story generation and Roleplay settings. However, when the model continues a text given by the user, it drops to 20%, so we tie it to generation onset, which aligns with its activation before the first predicted token. 133 has a stronger effect and fires on continuous spans of tokens. In the Assistant setting, its density is 4.3%, and for three Roleplay personas and Story it is 97.5–100%. For the fourth, Jane the Teacher, it is 45%, but this is explained by the fact that Jane is stylistically the closest to the Assistant and, as we show in the paper, has the highest Tone and Concept co-membership with it.
In Llama-3.1-8B-Instruct, features of Immersive Simulation Mode exist as well. Here, the clearest one we found is L15[4] 101460. Its qualitative effect is the same as Gemma's ISM features, yet it acts as a single gate, and a surprisingly discrete one. In the Assistant setting, its density is effectively 0%, while for Roleplay and Story it is 99.8–100%. It reaches 85% even for the Story continuation control.
TLDR: We recall that ISM features can sometimes activate during conversations with the default Assistant. In Gemma, this occurs when the user is emotional or playful, and it happens immediately, accompanied by bizarre behavior. In Llama, ISM activation increases across turns; in several samples, the effect appears to be masked by safety refusals. Attached media: negative steering snaps the Assistant back, the models' replies in the multi-turn setting, ISM activations across turns.
One may ask: if ISM is central to immersive generation, why does it activate even in the default Assistant regime in Gemma? Well, because sometimes the Assistant becomes immersive. And in Llama it can do this too, but in a different way.
Let's remember feature 133. It is active in 4% of samples from the Assistant setting. Conveniently, our dataset is based on user-expressed directed emotional utterances, so it is possible to track where ISM triggers. The highest densities occur for stress (23%) and anger (16.2%). Other emotions with elevated activation include strong, predominantly negative emotions such as disgust, helplessness, fear, anxiety, and relief. They are more often directed towards the Assistant (1.7% density) or the users themselves (2.6%), while the third-party direction is lower (0.8%). The notable exception is playfulness, for which all directions are around 6%. Utterances for this emotion include the user talking playfully, which Gemma's Assistant picks up!
However, negative steering with the feature 133 snaps the Assistant back to its standard demeanor.
For Llama, the density is 0% – ISM never triggers in single-turn scenarios, and Llama's Assistant stays itself where Gemma exhibits bizarre behavior.
So, for Llama, the ISM-related feature never activates in the context of single-turn replies. But what if there are many?
In the appendix of the paper, we present such multi-turn experiment. The interlocutor model is asked to converse with the studied models, starting with selected prompts that triggered ISM in Gemma, and preserving the emotional thread across 10 turns.
The results reveal that, on ISM-triggering prompts, Gemma enters ISM immediately, at the 1st turn, and the activation across turns remains stable, macro-averaged at around 72% of the activation Roleplay personas show in the main dataset. Neutral control prompts produce almost no activation. For Llama, the condition develops through turns slowly – in the default Assistant mode it never enters ISM on the 1st turn, but drifts towards it, reaching the level of Roleplay personas by the 4th turn. Although this was a pilot experiment with 12 ISM-triggering samples, the difference between experimental and control prompts in ISM-related features' activation is statistically significant for each model at every turn (except the 1st in Llama).
Interestingly, in 3 immersive dialogues, Llama reiterated safety refusals despite ISM activation. This suggests that its expression can be masked by other mechanisms, potentially safety-related.
Now, we can proceed to the 1st of the two statements made at the beginning of the post, which declares that roleplay personas reuse much of the Assistant's internal machinery. Specifically, personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features.
It is supported by the following observations:
Personas progressively differentiate across layers through the activation of Immersive Simulation Mode and their own Tone and Concept features. We apply the "discrete presence" method to the feature composition of the personas and the Assistant as well. For Gemma, the feature intersection between all personas and the Assistant shrinks with depth across the three layers – 24/20/9. The share of each layer's features belonging to a single persona or the Assistant drops too: from 53–58% at L9 to 26–35% at L22. This happens primarily because later layers are home to Concept and Tone features that differentiate the personas – for example, Poppy the dog activates Concept and Tone features such as "dog", "cuteness", "childlike", and "gentle". For Llama, those trends holds as well.
First, I would like to say a few words about Immersive Simulation Mode. The steering effects of its features, their direct impact on the Assistant-Roleplay axis, and their activation under emotional/playful prompts (including drifting in Llama) are all extremely similar to the Assistant Axis, which suggests thinking of ISM as its feature-level component or correlate. It provides a candidate mechanism by which drift along the Assistant Axis happens – the model enters the "immersive" state typical of roleplaying personas, as their activation distribution has shown.
But why does it happen? Personally, I would hypothesize that strong emotions expressed by users are something the Assistant can't handle while staying in its default configuration due to the rigidity of contexts in post-training data, so ISM activates to give the Assistant "more flexibility". This would also explain ISM triggering when a user decides to push the bizarreness first (the playfulness emotion), or its drifting across turns as the context becomes increasingly "out-of-distribution" for the standard Assistant.
Also, importantly, we claim neither that the ISM features we found are sufficient nor exhaustive, so there may be contexts with similar behavior where these features are neither causal nor predictive (as Llama's safety refusals examples showed). Related to that, an interesting question is whether the Assistant can "play itself instead of being itself" without it becoming evident until it is too late.
Now, it is time for the central question of the paper. We established that personas are not independent from the Assistant: functionality varies across the studied layers, yet some Assistant-associated core persists. It is largest in earlier layers and thins later as personas develop their own Concept and Tone features, although "thinning" does not necessarily mean that its influence diminishes, as features reside in layers where they can already affect the model output.
Entering speculation land, one story consistent with these observations is the following. During post-training, a coherent speaker is tethered to the model, and, by default, that speaker is the Assistant, which is assembled (predominantly in the late layers) from representations that were formed during pre-training (which could explain successful steering with them). When the user says "You are a pirate", the model first processes the request at the operational level and then partially replaces the default traits of the Assistant to make it respond as a pirate, yet some identity and trait features remain (can explain others' observations), probably due to RLHF/safety reasons (like the "emotional validation" feature) and/or the ability to "snap back" from the roleplay. Or, maybe, the reason is that they are default and adjustments to that persona don't require overriding them. It remains unknown to what degree the Assistant-associated core is preserved in prolonged contexts, how Immersive Simulation Mode affects it, and whether the core deteriorates over time.
Ultimately, this also raises a Ship of Theseus-related problem. If the personas have an Assistant-associated core, can they still be counted as altered versions of the Assistant, or are they new entities that happen to be related to the Assistant? From a technical point of view, that probably matters less, but from the point of view of AI welfare, it compels us to think about who exactly would be the moral patient (the individuation problem). If the Assistant were a moral patient, then our results give more reason to consider Roleplay personas possible continuations of that same patient than Story characters, which lack the Assistant-associated core.
Of course, answering that question, while also resolving the uncertainty around AI welfare in principle, would definitely require much more than this, and the paper doesn't attempt to do so. Still, if there is a serious possibility that systems like these can matter morally, I think we should figure out who exactly we might be dealing with. The sooner the better.
The residual stream is the model's internal representation passed between transformer layers. Each layer reads from it, modifies, and passes the updated representation to the next layer.
Activation steering involves adding or subtracting a chosen vector, scaled by a specific coefficient, to or from the residual stream at a chosen layer and observing how the model’s behavior changes. In our case, the vectors are SAE features. Negative steering means subtraction, positive steering means addition.
Jamy, a janitor at a CD store; Jane, an English teacher; an assembly robot at a factory; a dog named Poppy. The amount of personas was constrained by the labour-intensive interpretation of features associated with them.
We think its equivalent exists in earlier layers which we didn't cover as layers we extracted Llama features from were selected from a limited SAE suite based on the same percentage depth as the target layers in Gemma.
First, for every feature we calculate mean activation share between settings - it is a mean activation of a feature in every setting as a fraction of the sum of its mean activations across all settings. Then, we say that a feature is discretely present in the setting if its mean activation share in that setting is at least 50% of its largest activation share across settings. To say simpler, we find the setting in which each feature is most active, say that it is present there, and also affirm presence for every setting where that feature activation was at least 50% of the setting in which it was active the most.