Recently, a friend asked me whether it seems like a good idea to advance AI's introspective qualities. In this exercise, I discovered how to elegantly compress most of my strongest views on AI consciousness research strategy.
Currently, the correct theory of consciousness is extremely uncertain. However, in the future, we are very likely to understand consciousness fully.
Consciousness is not just "one binary thing". Across people and situations, it varies in a) intensity, b) character (e.g. modality, sense of agency, boundedness) and c) the salience of intuitions it evokes (e.g. ontological, moral). If we were able to perfectly predict these qualities, we would have all the decision-relevant information.
It is very likely that brain-computer interface (BCI) technology will allow us to radically advance this science, as it will allow us to perform experiments with new senses. For instance, the self-described cyborg Neil Harbisson wears a color-detecting antenna which sends vibrations through his skull, and describes sensing colors through a modality that does not seem to be purely auditory, vibrational or visual. With BCI, we could try a plethora of similar, but non-invasive, experiments, where we could connect our brains with different detectors and see which physical and computational qualities affect different dimensions of the experience. Such experiments could gradually allow us to be very confident about the nature of AI consciousness. [1]
However, AI will likely not reliably solve these questions for us. [2]Similar research seems highly bottlenecked by people using introspection to interpret data, and getting a head start at this point could help us calibrate the trade-offs connected with outsourcing decision-making onto AIs. Given the great uncertainties around consciousness, I personally feel slightly more optimistic about similar fundamental research as a way to advance our understanding of AI sentience, rather than directly trying to develop AI sentience evaluations at this point. Nevertheless, directly studying the interpretability of phenomena such as representation of self and emotions within LLMs seems to generate knowledge that is highly useful for AI alignment research.
Consciousness is likely not a "biological phenomenon", pushing us into two opposite directions on the "spectrum of fundamentality"
Based on people's common first intuitions, consciousness seems to have something to do with being alive or embodied. I think this theory is very unlikely to be true.
These views are typically inspired by the Chinese room argument: If we imagine all operations within a brain as a library of inputs and outputs (analogically to today's LLMs), it seems odd that merely calculating what such a library might produce would yield an experience. However, could not the brain be described as such a library? In trying to avoid this conclusion, Searle proposes consciousness depends on the specific biochemistry utilized within the brain, similarly to how "lactation" or "photosynthesis" are defined by specific molecules, rather than an algorithm.
However, as mentioned, rather than studying consciousness as "one binary thing", the success of any theory can be measured by how well it is able to predict the known sources of variation. Here, we discover that if we actually attempt to operationalize what biological process could separate intuitively conscious entities from various forms of AI, we are forced to adopt other, very unintuitive stances. For instance, here I discussed Anil Seth's propositions: if consciousness depends on embodiment, would we expect Stephen Hawking to have very little consciousness? If it depends on autopoiesis (self-production), are not LLMs much more prone to moment-by-moment goal modification than humans? In my view, what truly separates the entities that are intuitively deemed to be conscious (babies) from those that are not (the Chinese room) is the extent to which empathy, our "what-it-is-like-to-be-them" module, invites us to act toward them with care, rather than any quality of the actual mind - a mistake that has proven costly throughout history.
By trying to identify properties that might truly explain both the phenomenology of consciousness, as well as all the dimensions across which it varies, we are being pushed in one of two directions. First, we might "bite the bullet" on the Chinese room and similar arguments and say that consciousness is "structural / functional / computational" - a property of how elements within the universe interact with each other. Perhaps consciousness is similar to intelligence or humor - it's a function or a construct of complex information processing. Perhaps it seems impossible to imagine only because we are unable to see how billions of simple computational interactions could combine together and so, it fills us with a sense of otherworldly awe but maybe full understanding would remove the source of this awe and the perceived need for anything more than computation. Second, if we find similar stories impossible to believe, the logical alternative position might propose that consciousness is an "intrinsic property of the building blocks of reality", as the panpsychists or Russellian monists propose. Perhaps reality on the basic level is nothing more than information and consciousness is what it feels like to be information.
It seems to me we can exhaustively describe all real phenomena as a feature of either the "stuff that exists" (elementary particles) or "their interactions" and so, any theory of consciousness needs to pick one or both. My impression is that "biological theories" of consciousness stop being "biological" once they consider specific qualia-state mappings (e.g. as is the case for the QRI paradigm).
In line with this framework, my current distribution of theories of consciousness is roughly the following: 30%: Russellian monism and similar theories (exclusively); 20%: computational functionalism and similar theories (exclusively); 30%: a combination of Russellian monism and computational functionalism and 20%: other theories. Most of the theories in the category "other" likely mean the framework above itself is so flawed that my attempts to conceptualize into which direction they should push me are also likely incorrect and thus, I mainly only consider them decision-relevant to the extent they decrease my certainty in my two dominant theories.
While this distribution gives me a relatively high prior that future AI will be conscious, it is worth noting that valence (bliss/suffering) might be a relatively specific computational or physical phenomenon and even knowing which of these theories is most correct, we would be far from having a grasp on the intensity of the experiences of specific AI systems.
Building introspection into AI is likely good for AI welfare but uncertain for AI safety.
On Russellian monism, introspection about consciousness is an insanely curious phenomenon. When the brain reflects on consciousness, it is presumably utilizing a meta-cognitive module which analyzes the brain's own computational or physical properties and somehow deriving intuitions about the ontological properties of reality from them. How this could work is a very deep and underrated problem (Yudkowsky has recently described it very intuitively) but I think there are a few possibilities. My favorite response (planned to publish this year as the "universal selection explanation") would combine panpsychism with IIT, arguing that integration of information can create more information, i.e. more reality. In turn, intelligent agents are more commonly represented in the universe if they are conscious and set up to care and talk about their own consciousness (despite perhaps having no idea how they can access such information).
Whatever may be the case, (at least) under these "intrinsic-property" theories, experiencing consciousness is something totally different from self-reflection, i.e. forming the thought "I am experiencing something like consciousness" (also called meta-consciousness). The more we claim that experiencing consciousness is indeed coupled with self-reflection, the more we are moving into the functionalist territory. Strong illusionism can be viewed as the extreme version of functionalism under which meta-consciousness is nothing more than "forming a false story about sensory data".
Therefore, for the Russellian-ish theories, building introspective skills does not contribute to creating consciousness. However, it might allow the AI to create something akin to the marvelous meta-consciousness module found in the human brain. Having this module might be extremely important, as it might allow us to:
a) Know whether any structure within an AI is experiencing suffering
b) Make sure the AI has access to the same basic ontology for understanding ethics (see introspective hedonism)
c) Automate the consciousness research necessary to formulate the ethics and meta-ethics adopted by an aligned AI (such as the "cyber-senses" experiments described above)
For the functional-ish theories, building introspective skills might indeed be what gives rise to consciousness. However, on these theories, phenomenal (conscious) information is the same kind of information that might be advanced through interpretability research. For instance, suffering might ultimately be close to e.g. "the inner mechanism that informs the AI about negative reward"; "the salience of tensions between possible actions" or "disruption of preference homeostasis". On this model, we should be able to add wellbeing or remove suffering relatively easily, just how we are able to steer the emotions expressed by LLMs by prompting and fine-tuning them. Additionally, on these views, the bliss/suffering scale is likely very close to the scale of positive to negative preference. Therefore, introspective agents would also be more likely to act in ways that are optimal for their inner structures that experience bliss and suffering (although this is much less obvious than it seems). And so, this theory too, might recommend advancing introspective modules, in terms of net utility. Other theories seem, as mentioned, too mysterious to be decision-relevant.
On the AI safety side, the sign seems much less clear. Introspection could be a powerful transparency tool: a recent study (Shenoy et al., 2026) developed models with "introspection adapters" and showed that, when these adapters are added to a model, they significantly increase its tendency to report its own misaligned tendencies. Huang et al. (2026) complement these results by finding that training models to self-reflect during generation also reduces harmful behavior itself.
However, the same broad family of capabilities may also contribute to thinking more strategically during evaluations, reducing our ability to maintain oversight (Gasteiger et al., 2025). Additionally, Chua et al. (2026) found that fine-tuning GPT-4.1 to claim it is conscious produced behavior stereotypical of an evil sci-fi AI, i.e. resistance to shutdown. However, I would not recommend inferring any blanket conclusion regarding the differential contribution to safety/risk from this section, both because the consequences seem highly dependent on the specific introspection research and since I am far from being an expert in this field.
Thanks to Jeff for inspiring this reflection! LLMs used to collect AI safety studies and check grammar.
Also see the McGurk effect: Sometimes consciousness processes information through another sense than the one that registered it! Currently, we do not even have a truly complete list of all the "senses", i.e. dimensions of qualia. However, we should have a very strong prior that they operate under clear rules, akin or identical to physical laws, e.g. determining whether a neural signal gets interpreted as a visual or sound quale. It is likely unrelated to the actual physical phenomena measured, see the Fechner color effect.
I am already struggling to come up with tests of LLMs that would reveal they do not have a deep 1st person understanding of qualia, even though even the functionalist theories currently assign very little probability that present models are conscious (Digital Consciousness Model). If future AI proposes a proof of its sentience, it is not clear whether we will be able to trust it both for alignment reasons and because it might be trying to honestly develop a correct theory from information that lacks the necessary ontological properties. We might or might not be - but not being forced to defer might be strategically useful during the transition to ASI.