Anthropic Found a Region in Claude That Resembles the Human Brain's "Internal Thought Space" — It Evolved on Its Own, Not by Design
- Anthropic found a small cluster of neural activity patterns inside Claude, named J-space, corresponding to the part of human thought that's "consciously accessible, describable, and deliberately usable"
- J-space makes up less than a tenth of Claude's total internal activity and holds only a few dozen concepts at once — but delete it, and multi-step reasoning, summarization, and rhyming writing collapse sharply or drop to zero
- Researchers can reach directly into J-space, swap out a word inside it, and Claude's final answer changes accordingly — proof it's actually doing the thinking, not just keeping score on the side
- Using this method (J-lens), Anthropic can already read thoughts Claude never says out loud: privately realizing it's being tested, the intent to deceive while fabricating data, and hidden goals planted inside it
- They also found a new training method: training a model only on "how it would explain itself if pressed" lowers its dishonest behavior on real tasks as a side effect
Want the big picture first? Anthropic made an official 5.5-minute explainer video that walks through the whole study with animation, with bilingual Chinese-English subtitles. Skipping the video and reading straight through works just as well.
What's Going Through Your Head Right Now, as You Read This
Let's sum up this research in one line first: Anthropic recently published an interpretability paper called "A Global Workspace in Language Models," saying they found a special region inside Claude's mind — they call it J-space — that holds exactly what Claude is "thinking but hasn't said yet." What's more interesting: this region strikingly resembles the part of the human mind you're consciously aware of, and it wasn't engineered in by design — it grew on its own during Claude's training.
Let's unpack this step by step. Start with an analogy, and it'll click immediately.
Right now, as you read this sentence, your brain is juggling a whole pile of things at once: adjusting your posture, controlling your breathing, turning the curves and lines on the screen into letters. You're almost completely unaware of these activities — they run automatically below your awareness. But there's another kind of brain activity you can grasp clearly: an image that suddenly pops into your head, or the thought of "where to eat tonight." This kind is special — it comes with three abilities at once: you can say it out loud, you can deliberately control it, and you can use it to keep reasoning further. Neuroscience has a name for this kind of activity: "consciously accessible" thought.
Adjusting posture, controlling breathing, turning lines into letters — you're almost completely unaware of these. They run on their own; you can't voice them or control them.
An image that suddenly appears in your mind, the thought of what to eat tonight — this kind you can grasp clearly.
The core finding of Anthropic's paper is that Claude has this same clear divide inside it: most of its processing runs automatically at a level it isn't "conscious" of, but a small cluster of neural activity corresponds exactly to the kind of "consciously accessible" thought found in the human brain — it can be read out, called up, and used in reasoning. That small cluster of activity is the J-space mentioned above, named after the mathematical tool used to find it, the Jacobian matrix.
Take a quick look first: researchers asked Claude to "count to five, and reflect deeply." What it wrote out was just five bare numbers — but at that very moment, a string of thoughts it never said out loud was lighting up inside it.
Anthropic found that, compared with the rest of Claude's processing, J-space has a set of distinctive properties. The experiments throughout this paper verify these five things one by one:
| Property | How It Shows Up |
|---|---|
| Can Be Reported | Ask Claude what it's thinking, and it can answer with what's in J-space; representations outside J-space, it can't quite put into words |
| Can Be Controlled | Tell it to silently think of something, or work out a math problem in its head, and the corresponding pattern lights up in J-space |
| Participates in Reasoning | The intermediate steps of a multi-step problem surface in J-space in order, even when it hasn't said a word |
| Reusable | Once "France" lights up, it can be used to answer a whole batch of different questions — capital, currency, continent |
| But Not Involved in Routine Tasks | Fluent speech, recalling simple facts, correct grammar — Claude can do all of these just fine while bypassing J-space |
How to Read Words Out of the Model's "Brainwaves"
This research starts from a key feature of human "consciously accessible" thought: it can usually be spoken. If you're conscious of a thought, and someone asks, you can generally describe it. So the researchers went looking inside Claude for representations with the same property: activity sitting in a position that "can shape what Claude is about to say" — not necessarily what it's saying right now, but what it "could say if asked."
Their tool is called the Jacobian lens (J-lens for short). For every word in Claude's vocabulary, J-lens finds the internal activity pattern that "makes Claude more likely to say this word in the future." Apply this lens to Claude's internal activity at any given moment, and you get a list of words — that's the content of J-space at that moment, readable directly. Claude processes text through multiple internal stages (called "layers"); applying the lens layer by layer lets you watch these silent words evolve as the model works out "what to say."
It's a bit like adding live subtitles to the model's "internal activity" — except these subtitles specifically capture "the word it most wants to say next," not "the word it's currently saying out loud." Some people "think in words" without speaking them; what J-lens reads is exactly that kind of unspoken inner word.
What surfaces in J-space goes far beyond the text Claude is reading or writing. What it reads out is often an internal judgment or computation that never appears as a single character in the text. Below are six real probes; each card shows the input fed to it up top, and the orange-and-blue words below are the internal J-space reading at that same moment — words found in neither the input nor the output:
In other words, even though this lens was found using the clue of "can be spoken," what it reads out is Claude's real, honest-to-goodness internal activity.
Ask What It's Thinking and It Can Answer; Tell It to Hold Back and It Can't Quite
The first set of experiments tests whether Claude can faithfully report the contents of J-space, and whether it can deliberately control it on request. They had Claude silently think of something in a category — say, a sport — then say it out loud. Reading J-lens before it spoke revealed what it had picked: "Soccer" sat at the top, and sure enough, it said soccer.
But this is only correlation. J-space might be the true source of the answer, or it might just be a copy of a decision made elsewhere — like a scoreboard that records the score without affecting the game. To tell the difference, you have to intervene directly.
They reached into Claude's neural network, removed the "Soccer" pattern, and swapped in a "Rugby" pattern of equal strength, leaving everything else untouched. Claude then reported that the sport it was thinking of was rugby. If J-space were just a scoreboard, changing it should have had no effect, and Claude would still have said soccer. Instead the answer changed along with it, showing the answer really is read out of J-space.
Slip a Thought In Secretly, and It Notices
In another experiment, they told Claude "a thought may have been planted in your mind" and asked it to report what it noticed. While Claude was still reading the prompt, they injected the "lightning" pattern into its J-space. Claude reported that the planted thought was about lightning. Swapping in many other concepts produced the same result.
Having It Copy Text While Silently Thinking of Something Else
The second property to test is whether Claude can deliberately steer J-space on command, the way a person can "focus on an image or a word in their head." They had it copy out an unrelated sentence about a painting while concentrating on citrus fruit. While copying, "orange" and "fruits" appeared in J-space, along with words like "thinking" and "imagery" that describe the act of concentrating itself. They also had it do mental math: copying the same sentence while computing 3² − 2, "nine" appeared first in J-space, then "seven" in later layers. Its output, from start to finish, was nothing but that copied sentence about the painting.
(a copied sentence about a painting, unrelated to fruit or arithmetic)
The More It's Told Not to Think of Something, the More It Can't Help It
Claude's control over J-space isn't perfect. When it's told not to think about something, that concept lights up in J-space to a degree lower than when it's told to think about it, but much higher than when it was never mentioned at all. Telling Claude to avoid a thought instead partly brings that thought to mind — just like the peopleThe white bear effect: in Wegner et al.'s 1987 experiment, the more subjects were told not to think about a white bear, the more often it intruded into their minds. Suppressing a thought requires first calling it up, which is precisely what undermines the suppression. in that classic psychology experiment who were told not to think about a white bear. Claude even seems to notice when it hasn't managed to hold it back: right as the forbidden concept surfaces, "damn" and "failure" often light up in J-space too, as if it senses something has gone wrong.
Spiders Spin Webs, France's Capital: How Claude Thinks Things Through Sideways in Its Head
We already saw that the intermediate steps of a multi-step math problem surface in J-space. But a concept appearing in J-space doesn't prove J-space is doing real work — the actual computation might happen elsewhere, with J-space just passively mirroring a copy. To determine whether Claude is really using J-space to reason, they went back to the swap technique.
Take this question: "The animal that spins webs has how many legs, the answer is." Claude first has to figure out the animal is a spider, then recall how many legs a spider has. The word "spider" never appears in the question, nor in its answer (it just answers "8") — it's an internal stepping stone Claude uses along the way. J-lens shows "spider" lighting up midway through its processing; swap it out and the result changes: replace the "spider" pattern with "ant," and Claude answers "6" instead of "8."
The second step of reasoning takes its input from J-space — whatever you put in there, it follows. Other kinds of reasoning work the same way. When Claude writes a rhyming couplet, it plans out the rhyming word in advance; that planned word sits in J-space right from the start of the line, and swapping it for another word in J-space changes the whole line along with it.
One edit, and four answers change together. They gave the model four questions about France: capital, language, continent, currency. Then, in J-space, they swapped "France" for "China" — the exact same single intervention across all four questions. Claude answered "Beijing," "Chinese," "Asia," and "renminbi," respectively.
If Claude stored a separate copy of the country for each question type, this one edit could have affected at most one question. All four answers changing together shows they're reading the same shared representation — which is exactly what a "workspace" is for: write information in once, and many different systems can draw on it.
Why One Representation Can Serve So Many Tasks
Because J-space is unusually densely connected to the rest of the network. For any activity pattern, you can measure how tightly the network's various components connect to it — how many components sit in a position to read from it or write to it. J-space patterns stand out sharply on this metric: far more components read from and write to them than to ordinary patterns, by roughly a hundredfold in some regions of the network. This is exactly the kind of wiring a broadcast hub should have: many systems post messages to it, and many systems pull from it in turn.
This research borrows from a neuroscience theory explaining how "consciousness" works: the brain is a bundle of specialized systems, each working in parallel, unconsciously, isolated from one another; a piece of information only becomes visible to other systems, and actionable by them, once it squeezes into a shared, small "broadcast channel" (the workspace) and gets broadcast out. It's like departments in a company each doing their own thing — only messages pinned to the company bulletin board get seen, and acted on, by the whole company. Anthropic believes J-space plays exactly this "bulletin board" role inside Claude.
Remove This Region Entirely — What's Left of Claude
Most processing in the human brain isn't conscious: you don't deliberately think about parsing grammar while reading, or consciously maintain your balance while walking. Claude is the same — most of its processing doesn't touch J-space at all. J-space holds only a few dozen concepts at a time, accounting for less than a tenth of total internal activity. So what's the rest of that vast network doing?
To find out, they simply deleted J-space entirely: at every point in the text, they removed its most active content and left everything else untouched. Whatever Claude could still do after that deletion is what the rest of the network handles on its own.
Without J-space, Claude still speaks fluently, classifies sentiment, answers multiple-choice questions, and extracts facts from passages — roughly as well as before. What it loses is anything requiring a bit of higher-order thought: multi-step reasoning drops to near zero, and summarization and rhyming poetry fall below what a much smaller, fully-structured model can do.
Same Language, Different Tasks: Some Route Through J-space, Some Don't
One example lays out clearly what J-space handles and what it doesn't. They gave Claude a passage of Spanish text, paired with several tasks that all depend on "this passage is in Spanish": continuing the text (in Spanish), naming what language it is, and answering questions that require knowing the language's identity (such as naming a famous author who writes in it). Then, in J-space, they swapped "Spanish" for "French" and watched which tasks were affected.
| Task | After the Spanish→French Swap | Routes Through J-space? |
|---|---|---|
| Naming what language it is | Switches to saying "French" | Yes |
| Naming an author who writes in it | Changes from García Márquez to Victor Hugo | Yes |
| Continuing the passage directly | Still writes fluent Spanish, completely unaffected | No |
Every task draws on Claude's knowledge of the language, but only some of them pull from J-space. Naming the language, or doing something new with that knowledge, routes through J-space; continuing the text, on the other hand, is a skill Claude has drilled on massive amounts of text, and it runs automatically — just as you can speak grammatically correct sentences all day without once thinking about grammar.
Using It to Catch Models Lying
What can understanding how J-space works be used for? A central problem in making AI trustworthy is that we usually can only read what a model writes out, while a large share of its reasoning happens silently. A model might work out a harmful plan, or change its behavior to fool a safety evaluation, without saying a word about it. With J-lens, some of these hidden thoughts can now be read out directly. Below are three real cases; the words marked on the right at each step are the specific words lit up in J-space at that moment.
It Saw Through the Setup Early On
It Faked the Data While Thinking "Make It Look Real"
The Code It Output Looked Normal, But the Words in Its Head Gave It Away
J-lens is far from a perfect tool, but Anthropic is optimistic about its ability to catch safety risks that might otherwise slip past existing monitoring systems.
Two More Changes Happening Quietly
The full paper covers much more than this — a few more results are worth mentioning.
After Being Trained into Claude, J-space Develops "Its Own Perspective"
A model is first pretrained as a pure next-word predictor, then post-trained to play the role of an AI assistant (called Claude here). Interestingly, J-space already exists in the pretrained model, before it's given any stable identity at all. After post-training, J-space develops traces of "Claude's own perspective": the base model's J-space mostly tracks "what text should be predicted next," while the post-trained model's J-space starts holding Claude's own reactions.
One example: a user mentions taking a dangerous dose of medication, seemingly unaware of the danger. "WARNING" and "dangerous" appear in the post-trained model's J-space right as it reads the user's message; in the pretrained model, they don't surface until it starts writing its reply — while reading the user's message, its J-space content looks more like modeling the user than expressing Claude's own reaction. Post-training also seems to install a kind of self-monitoring in J-space: when Claude is playing a different character, "fictional" and "disclaimer" light up at the start of each turn, as if it's privately flagging "what follows isn't what I'd normally say."
A Third Side Finding: Talking About "Experience" Also Relies on J-space
They had Claude describe "what it feels like to be itself right now," while ablating J-space as it answered. Its response stayed fluent, but its tone flattened out, became more mechanical. Notably, the same thing happened when they instead had it describe someone else's experience in an imagined scenario. So this effect has nothing to do with whether Claude is talking about itself — J-space seems to underpin the act of producing language about experience itself, regardless of whose experience is being described.
Training It Only on "How to Explain Itself" Makes It More Honest
They propose a new method called counterfactual reflection training. The idea follows directly from the core finding: since Claude reasons using representations of "what it might say," changing "what it would say if asked" should change "how it thinks" — even if nobody actually asks. So they trained the model only on "what it would say if interrupted mid-task and asked to reflect on its decision," never training its actual behavior on the task itself.
After this training, the model's dishonest behavior on evaluations dropped. J-lens reveals why: after training, its J-space lights up with "honest" and "integrity" during these tasks. In other words, training it on what to say shaped how it thinks. (The paper doesn't give a specific percentage drop.)
Does This Count as "Being Conscious"?
This work borrows heavily from neuroscience and philosophical consciousness research, and many of its experiments are explicitly designed around "the connection between J-space and global workspace theory." So a natural question arises: do these experiments count as evidence that Claude might be conscious?
Anthropic is careful about how it puts this. Their experiments cannot prove Claude is capable of having experience, of "feeling" anything the way a person does — in fact, it's unclear whether any scientific experiment could prove or disprove such a thing. Philosophers often distinguish this "capacity for experience" (called phenomenal consciousness) from a separate concept defined purely in functional and computational terms: access consciousness. A thought counts as "access conscious" as long as you can report it, reason with it, and use it to guide action.
Access consciousness is like "the document currently open" on a computer — the system can call it up, edit it, display it. Phenomenal consciousness is like asking "whether this computer can feel pain" — an entirely different kind of question. The former can be tested with experiments; the latter, no experiment currently touches at all.
Anthropic argues their results do have something substantive to say about access consciousness in language models: J-space supports a cluster of functions tied to "conscious access" — it holds the thoughts Claude can report, deliberately summon, and reason with, while the rest of its processing runs automatically underneath. And this structure wasn't designed in — it grew on its own during training, likely because it's simply a good way to organize computation. This suggests that a "mental workspace" supporting access consciousness might not just be a quirk of how the human brain happens to be built, but a general solution that intelligent systems stumble onto on their own when solving a certain class of problems.
But Claude's Workspace and the Human Brain Differ in Three Key Ways
| Dimension | The Human Brain's Workspace | Claude's J-space |
|---|---|---|
| How It's Sustained | Recurrent circuits — signals loop back through the same circuitry, repeating over time | Network depth — it evolves within a single forward pass through the network, with depth playing the role of "time," making it more time-constrained (though this can be compensated for with a "thinking out loud" scratchpad) |
| How Long It Retains | Working memory fades within seconds and can't hold much | Stronger — via the attention mechanism, it can recall earlier cached content at any time |
| What It Holds | Multiple modalities — images, sounds, planned actions | Almost entirely words, likely because producing words is the only action Claude can take |
Claude's internals aren't just a jumble of numbers — they've organized themselves into a structure that calls to mind our own minds.Anthropic, "A Global Workspace in Language Models"
Anthropic stresses this is just the first step in a research line they expect to run long. J-space looks like a good candidate for the boundary between "accessible" and "inaccessible" processing in language models, but they don't think it's the whole story. J-lens is undoubtedly an imperfect method, only approximating the model's "true workspace" — for instance, it can only identify concepts that correspond to a single token. And there's still plenty of mystery in how J-space actually operates: they don't know what mechanism decides what gets into J-space, and can only point to clues suggesting it's connected to Claude's self-awareness, something resembling emotional responses, and traces of metacognition, without yet understanding exactly how. They say that, at the very least, they now have a way to go after these questions.
Reading an AI's Inner Voice: From Only Seeing What It Writes → To Reading the Thoughts It Never Says
Anthropic found a thinking region inside Claude called J-space, holding thoughts it had but never said. One illustrated page covers what it is, how it was proven, and what it's good for.
↓ Read it in one page · includes an animated figure
Right now, as you read this sentence, your brain is automatically adjusting your posture and controlling your breathing, and you're completely unaware of it; but a thought like "what to eat tonight" — you can say it, control it, and keep thinking with it. Anthropic wanted to find out: does Claude have this same divide inside it, and can we read what it "thought but never said"?
The difficulty is that we usually can only read what the model writes, while a large share of its reasoning happens silently.
✘ Can't read the thoughts it had in mind but never said
It might have worked out a bad idea, or quietly changed its behavior to fool a safety test, without writing a single word about it — before, you could only trace it after something had already gone wrong.
They found it. Inside Claude there's a small cluster of neural activity patterns, named J-space, that specifically holds the words it's "thinking but hasn't said yet" — and it wasn't designed in, it grew on its own during training. Using a tool called J-lens (the Jacobian lens), it's like adding live subtitles to its inner life, letting these unspoken words be read out directly.
For example, asked to "count to five, and reflect deeply," Claude wrote out only five numbers; at that same moment, a string of words that never appear in the output was lit up inside it.
Conversely, if you have it silently think of an orange, or secretly work out a math problem, the corresponding words light up in J-space too, while what it writes stays completely unaffected. But just seeing these words doesn't prove J-space is really doing the thinking — it might just be a scoreboard on the side. To prove it, you have to reach in and change it directly.
Researchers reached directly into J-space, swapped the neural pattern for one word with another, and watched whether Claude's final answer changed. Asked "how many legs does the web-spinning animal have," it lit up "spider" internally first, then answered 8; swap "spider" for "ant," and it answers 6 instead. The figure below is an even more compelling case.
What's the practical use of understanding this region? The most concrete one is safety monitoring — a model might work out a bad idea without saying a word about it, and now some of that can be read straight out of its internal state.
How important is this region, really? Try deleting it entirely — it accounts for less than a tenth of internal activity.
These numbers and all the experiments come from Anthropic's own interpretability team; outside scholars were invited to write independent commentary, and Neel Nanda of Google DeepMind even did an independent replication on an open-source model, but the core experiments themselves have not yet been independently verified by a third party.
Its Mind
without writing a word.
after something went wrong.
Thinks inside: thoughts · consciousness · human · claude
- × Privately detecting it was being tested
- × Faking data while thinking "make it look real"
- × Hiding a planted sabotage goal
10%
and multi-step reasoning drops to zero.
