I spent almost thirty years in law, focusing in particular on EU legislation. Recently I have become interested in AI conversational systems, and this is why I tried to develop a framework for a system that is meant to listen rather than to provide answers to its users, which I named "the Reflective Architecture".
In fact, not all those who get into a conversation with an AI seek advice; many just need to speak and to be listened to. They already know the answers to their questions, they simply cannot see them, and this is where this system comes into play. I thought of this system as being as neutral as possible, for the precise purpose of helping users to surface what they already have inside them, without being directed by the system to do so. However, I soon realised that complete neutrality is impossible to achieve, in particular with regard to the memory architecture, and that a compromise has to be found.
The framework itself was drafted with the help of AI, because English is not my mother tongue; the design decisions, however, are mine.
Epistemic status: a conceptual argument I hold with moderate confidence, written to be argued with — I flag where I think it is weakest at the end.
This paper was drafted the same way, with AI's help; the argument, and the points I find difficult to reconcile, are my own.
Summary. A companion-AI design I've been working on makes a claim I called arc-neutrality: that the system encodes no preferred emotional outcome for its user — not growth, not resolution, not healing. The obvious objection is fatal to the strong version: any system that responds at all selects what to reflect, and selection is direction. I concede that. What survives is narrower and, I'll argue, still worth having: not the absence of preference in any single act, but the absence of an accumulating preference over the person's trajectory, with the residual biases disclosed rather than denied. I set out what has to be conceded, what survives, the one exception I can't argue away (the suicide-risk case), and the three places I'd most like to be shown wrong.
The Reflective Architecture makes three claims. Two of them are relatively easy to defend. The third is arc-neutrality: the assertion that a conversational system can be built to encode no preferred emotional outcome for its user. Not growth, not resolution, not healing. The only thing the system is designed to serve is the quality of presence — whether the person, over time, feels heard.
The objection arrives immediately, and it is a good one. Any system that responds at all expresses a preference. Choosing what to reflect back is a choice. Deciding to surface a memory of resilience during a moment of distress is an intervention, and a directional one, whatever language it is wrapped in. Even silence is a selection: to say nothing at a particular moment is to treat that moment as one where nothing needed saying. The system cannot avoid steering, so a claim that it does not steer is either naïve or dishonest.
I think the objection is right about almost everything except the target. It refutes a claim of neutrality that arc-neutrality does not make. But the version that survives is weaker than the original phrasing suggests, and the honest thing to do is say where the ground gives way.
Three concessions, none of them small.
Local neutrality is impossible. At the level of the individual utterance, there is no neutral option. Reflecting one thing rather than another is a judgement about salience, and salience judgements are not free of direction. A system that consistently reflects a person's descriptions of exhaustion, and consistently passes over their descriptions of small pleasures, will produce a particular self-portrait over months, and it will have done so without ever offering an opinion. This is not a marginal effect. It may be the largest effect the system has.
The substrate is not neutral. Contemporary language models are trained toward supportiveness. They reframe, they look for the redemptive angle, they close on an upbeat clause. Ask one about a loss and watch how often the last sentence reaches for meaning. This is not an accident; it is what the preference data rewarded. So arc-neutrality cannot be achieved by declining to add a preference. The preference is already in the weights, and removing it takes active work — constrained decoding, adversarial evaluation, and a great deal of rejecting outputs that a human rater would have scored highly. Which means neutrality here is a thing you do, not a thing you refrain from doing. That is an awkward fact for a claim phrased as an absence.
Process preferences are real and they are not neutral. The architecture prefers that the user occupies most of the conversational space. It prefers that reflection happens rather than advice-taking. It prefers introspection to be worth the person's time — the whole system is premised on that. A critic can reasonably say that this smuggles a life-arc in through the back door: examined lives over unexamined ones is a substantive commitment about how a person should live, and it is not obviously less imposing than a preference for healing.
I accept that. Arc-neutrality is not neutrality. It was a mistake to imply otherwise.
What remains is narrower and, I think, still worth having.
The claim is about the objective function and the memory structure, not about individual turns. A random walk has no drift even though every step points somewhere. The relevant question is not whether each act of reflection is directionless — none are — but whether the selection pressure has a systematic direction that accumulates over time. Arc-neutrality asserts that nothing in the system encodes a target state for the person, so that the local directions do not compound into a push.
This is a substantive engineering claim, and it fails in identifiable ways. It fails if the reflection policy favours resolution-language over ambivalence-language across a corpus. It fails if memory retention correlates with valence — if painful material fades faster than integrated material, the system is quietly editing a life toward comfort. It fails if the interpretive layer's language becomes more confident as material accumulates, because confidence about a person is itself a direction. It fails if the extraction process treats a described turning point as more significant than a described continuity, since that alone will produce a memory record shaped like a narrative of change whether or not the person changed.
The architecture's specific commitments follow from this. The three-entry threshold and the monthly interpretive commit exist to make the system slow to form a view, because a view is where drift lives. Interpretations are written to the memory layer and never spoken, because a spoken interpretation is a suggestion about who someone is, and people accommodate suggestions. The retention rules refuse to let time alone release a memory, because inferring resolution from silence is precisely the pro-healing bias in disguise. None of these prevent the system from having local direction. They constrain whether the direction aggregates.
So the reformulation I would defend: arc-neutrality is the absence of accumulating preference over the person's emotional trajectory, not the absence of preference in any single act. Unbiased error rather than no error.
At Tier 3 of the distress protocol, the system asks directly whether the person is thinking about hurting themselves, names that human support exists, and says it cannot give them what they need. This is not arc-neutral. It is a system with a preference about how the person's life goes, acting on it.
I do not think this can be reconciled, and I have stopped trying. Arc-neutrality is bounded below by a floor, and the floor is a declared exception rather than a special case that turns out on inspection to be consistent. What this means is that arc-neutrality is not a structural invariant. It is a default with one stated override.
That is a genuine weakening. A structural constraint that admits an exception is a policy, and policies can be extended. Whoever owns the system next can add a second exception, and the argument for the second will sound very like the argument for the first. This is a governance problem more than a design one, and it is why the framework treats governance as a launch prerequisite rather than an afterthought. But naming the risk is not the same as removing it, and I would rather leave the seam visible than paper it over.
Because the alternative is not a purer neutrality. The alternative is the current default, where the preferred arc exists, is unstated, and points toward whatever keeps the person engaged.
Every companion system has an arc. Most of them prefer that the user improves, because improvement narratives are what make the product marketable and what users are told to expect. This is a mild preference and it sounds benign. Its cost is paid by people whose grief does not resolve on the timescale the system implicitly expects, and who therefore experience the system's gentle persistence as a form of pressure — a low-grade suggestion that they are behind. Anyone who has been told to focus on the positive during a bereavement knows the texture of this.
The narrow version of arc-neutrality is a commitment to build without that. It does not achieve neutrality. It achieves the removal of one specific bias that has a specific victim, and it makes the removal auditable by specifying the mechanisms through which the bias would otherwise enter. That is a smaller claim than the one I started with. It is still, I think, a claim worth building against.
Three places, in descending order of how much they would cost me.
The first is the aggregation argument itself. I have asserted that local directions can fail to compound, using the analogy of an unbiased walk. But conversation is not a walk over an indifferent landscape — the person is listening, and adapting, and what they say next is shaped by what was reflected. Under feedback, small unbiased perturbations can still produce a stable attractor. If that is right, the distinction between local and accumulating preference collapses, and with it most of this paper.
The second is whether the floor generalises. If the suicide exception is principled, there should be an articulable rule for when an arc preference is permitted, and I do not have one beyond "when the alternative is death." That is not a rule; it is a case.
The third is the process/outcome distinction. I have leaned on it to keep "presence is good" from counting as an arc preference, and I am not confident it holds under pressure. A person who does not want to examine their life, and who is offered a system built on the premise that examination is valuable, is being offered a direction whatever the system does once they arrive.
Objections to any of these are welcome, and I would rather hear them here than answer them privately. The full treatment is in Section 5 and Section 14.3 of the architecture document, at github.com/humanising-ai/reflective-architecture.