Swiss Unic AG
DEENStart a conversation

What Claude processes but does not say: J-Space and the question of consciousness

Visualisation of internal AI representations in J-Space

Language models process more internally than they reveal in chat. Anthropic researchers have identified a small collection of internal activation patterns in their language model Claude that they call “J-Space”. This raises a fundamental question: Is Claude conscious?

Unspoken concepts: What is J-Space?

Have you ever wondered how a chatbot handles a question internally before formulating its answer? In Claude, certain internal activation patterns can be observed that are associated with concepts even when the corresponding words never appear in the final text. This functional separation is precisely what researchers at the AI company Anthropic investigated. On 6 July 2026, they published the interpretability study “A global workspace in language models”. In it, they describe a small collection of internal representations they call “J-Space”. The name derives from “Jacobian”, a mathematical term central to the analytical method used – the so-called “Jacobian Lens”, or “J-Lens” for short.

According to the researchers, J-Space was not explicitly programmed as a separate module. Instead, this structure emerged during training. The study shows that Claude can access the contents represented there comparatively well, deliberately activate them to a certain extent, and use them for multi-step tasks. The term “workspace” is an analogy to Global Workspace Theory in consciousness research. In this theory, information becomes consciously accessible when it enters a limited workspace available to different processes.

J-Space differs from human memory in three respects:

Not a conventional memory store: J-Space refers only to activation patterns within the model's processing.

Invisible to users, but analysable: These patterns do not appear in an ordinary chat. Using specialised interpretability methods, however, researchers can approximately examine and deliberately alter them.

Not the same as chain-of-thought: J-Space is not a written-out sequence of reasoning.

A broadcast hub in the silicon brain

J-Space is not directly accessible to ordinary users. Experiments by Anthropic researchers nevertheless show that it can be influenced. One example makes this particularly vivid: if Claude is asked to silently think about citrus fruit while copying an unrelated sentence, the J-Lens analysis reveals, among other things, the representation “orange”. The concept is therefore detectable internally without appearing in the generated text. Similar experiments show that intermediate representations in multi-step tasks can also become visible in J-Space.

These findings are reminiscent of the Global Workspace Theory mentioned above. It originated with cognitive scientist Bernard Baars and was later developed further by Stanislas Dehaene and other neuroscientists. The basic idea is that many specialised systems process information in parallel without those contents being globally accessible. Information becomes functionally consciously accessible when it enters a limited shared workspace from which it can be broadcast to numerous other processes.

Anthropic examined five functional properties of such a workspace system. In J-Space they can be described as follows:

  1. Reportability: When asked, Claude can name contents that correspond to representations identified in J-Space.
  2. Flexible generalisation: The same representation can be used by different downstream computations across different tasks.
  3. Directed control: An instruction can bring a concept into J-Space and keep it active across several processing steps.
  4. Internal inference: Unspoken intermediate steps can appear in J-Space and contribute causally to the result. If such a representation is deliberately replaced, the answer can change as well.
  5. Selectivity: J-Space comprises only a small part of the model's overall activity. Many routines, including parts of language processing, apparently do not require this structure.

Of white bears and pink elephants

A particularly vivid finding concerns negative instructions. What happens when Claude is told not to think about something specific? As with the human “white bear effect”, the prohibition itself can activate the very concept that is meant to be suppressed. When suppression failed, researchers also observed representations in Claude associated with frustration or failure, such as “damn” or “failure”. Even language models, it seems, cannot simply avoid thinking about the proverbial pink elephant.

There is, however, an important difference between AI and humans in how information is handled. Human working memory is limited and can quickly lose information unless it is actively maintained. Claude, by contrast, can refer back to text much earlier in its respective context window.

What Anthropic's J-Space study reveals about AI safety

What happens when an AI model recognises that it is being tested? Does it change its behaviour? And can strategic or deceptive considerations be detected before they appear in the visible output? Anthropic's J-Space study offers a cautious but important answer: under controlled experimental conditions, internal representations can be identified that are associated with evaluation awareness, strategic considerations and deceptive intent. This makes the J-Lens method particularly interesting for AI safety research.

In alignment evaluations, the researchers found indications of strategic concepts such as “leverage” and “manipulation”. In situations where the model appeared to recognise that it was in an artificial evaluation or role-play scenario, representations such as “fake” and “fictional” also appeared. These results suggest that a model may internally assess a testing situation differently from what its visible answer reveals. Over the longer term, Anthropic's “Counterfactual Reflection Training” may prove particularly promising as an approach for deliberately influencing safety-relevant behaviour in language models.

Access consciousness: yes. Subjective experience: unresolved.

So how close does this bring Claude to consciousness? The scientifically careful answer is: it depends on what we mean by “consciousness”. Anthropic distinguishes here between “Access Consciousness” – functional accessibility – and “Phenomenal Consciousness”, meaning subjective experience.

The functional side includes capabilities such as deliberately activating, reporting and suppressing contents, and using them for multi-step reasoning – flexibly across different tasks. For these properties, the study finds interesting parallels with a global workspace in the models examined. But it explicitly does not claim that Claude therefore has feelings, a subjective perspective or human consciousness.

The most important conclusion for companies is therefore not that every AI system will soon be completely transparent. Rather, it is this: safety evaluations should not examine visible output alone. Analyses of internal representations could become an additional evaluation channel in future – alongside access restrictions, behavioural testing, red teaming, monitoring and human oversight.

J-Space is a promising but still limited diagnostic instrument. That limitation is precisely what makes it significant for AI safety policy: we are beginning to examine not only what a chatbot “says”, but also what it appears to be “thinking” while doing so. What is your view?