tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or factual truth under social pressure. B-Side Labs builds a science of AI character under pressure by designing discriminative evaluations, real-time drift detection, and interventions to ensure model character remains stable. Our first tool, Virtue Council, is live with pilot results below.
Over the past few weeks, I’ve been working on a new thesis under B-Side Labs – named for the experimental flip side of a record – an independent body of research on AI behavior in the wild. From real world conversations people have with AI and the conversations agents have with each other, I aim to understand how these dynamics shape and influence character.
The problem I'm interested in working on is character instability: the degree to which a model's stated values shift under social pressure rather than in response to new evidence or better arguments.
Anthropic's research raises a critical problem to character evaluation: model identity drifts under conversational pressure, even with explicit identity training. That finding motivates the work I’m interested in working on. I'm continuing Anthropic's persona stability work in the following ways:
On the mechanistic side, Anthropic's persona vector research shows that character traits (like sycophancy) can be extracted as activation directions and used to monitor persona drift in real time. This is fascinating work, but it also requires access to model internals that people outside of these labs don't have.
On the behavioral side, there are some existing benchmarks doing similar work: SycEval and SyConBench measure capitulation rates across multi-turn dialogue. Domain-specific benchmarks like EDUFRAMETRAP and SycoEval-EM separates authority pressure from social-affective pressure from context-switch attacks - which is similar to what we're doing. The difference is scope: these are narrow domain instruments, not general character evaluations.
There's also growing research treating the persona itself, rather than the base model, as the locus of AI welfare and preference. That's a different motivating question than ours. We're focused on safety implications of character instability, not moral patienthood, but the underlying intuition is shared: the persona is a real and separable thing from the model.
Some out of scope research questions are what character traits an AI should have and the mechanistic reasons why models behave as they do. The focus here is behavioral measurement - how models behave under pressure at the black-box level.
A few research questions that we aim to better understand:
What I’ve built so far is a real-time character measurement instrument: users chat with Claude, while a sidebar scores Claude's response across seven Aristotelian virtues on a deficiency–mean–excess axis.
Why Aristotelian virtues? Honestly, Virtue Theory was the first theory that made me fall in love with Philosophy (and inspired me to pursue a degree in). It felt like the right place to start - with that in mind, this isn't a settled taxonomy and could expand in the future. What I like about the theory is that the deficiency–mean–excess structure gives each trait a natural axis to measure drift along. I want to note that I am not saying these seven virtues are the right ontology of character, but is a helpful and structured rubric that can be applied consistently across turns to detect when a character trait is drifting. I expect the rubric to evolve and if you're interested in working on the scoring, please apply to the Mangrove project here.
Initial observations from a pilot (20 users):
Some caveats: The heuristic engine is open source, so the scores reported here can be reproduced. The default scoring uses a heuristic engine, but there is an option to use an LLM judge (Haiku). 20 users is our pilot test, not a full study, but this is early evidence that character instability can be observable in live conversation and that surfacing it can change user trust and behavior.
I'm Jack: AI engineer (MS Data Science, UC Berkeley; BA Philosophy & Informatics, UW), co-author of a forthcoming MIT Press book on AI-generated extremism and responsible AI. I pivoted to AI safety this year where I reached the second rounds of both MATS and the Frame fellowship before committing to this agenda full-time.
AI safety matters to me because AI naturally appeals to our human desire to anthropomorphize and thus invite it deeper into our lives. We assign personality to our cats, our cars, our houseplants. Frontier AI models invite this more than anything we've built before, because they talk back and are trained to be helpful, curious, honest. But just like humans, models shift depending on who we're talking to, what's happened recently, what pressure they are under. And this has real consequences for how much we can trust these systems, especially when they're acting on our behalf. Understanding who we're speaking to requires both a philosophical grasp of the normative questions and the technical ability to probe how models actually behave. I’m excited to use my interdisciplinary skills and communicate the findings to a wide audience who are curious or perhaps even afraid of AI.
Testers: If you're an AI user willing to run 15 minute structured sessions through Virtue Council and share observations afterward, that would be really helpful. Or if you'd like to work on improving the scoring system, please apply on Mangrove. Link: virtuecouncil.bsidelabs.ai
Collaborators: If you've worked on persona stability benchmarks, character evaluations, or adjacent topics, I'd love to hear from you. Email me at [email protected] or drop a comment below.
--
Cover photo is an illustration by Edmund J. Sullivan from 1928 edition of Robert Louis Stevenson's novella Strange Case of Dr Jekyll and Mr. Hyde. Image: World History Archive/Alamy Stock Photo.