cat mcgee
← back
July 8, 2026

The race to understanding

The real deadline is not when AI becomes too capable to control. It's when AI understands us better than we understand it.

When we talk about AI safety, we often frame it as a race: we must figure out alignment before the machines become too powerful to control. I don't think this is the right race.

The real deadline is not when AI becomes too capable to control. It's when AI understands us better than we understand it.

We're the tortoise and we're helping the rabbit

For some reason, nobody is treating this as a deadline. In fact, they're actually incentivising the other side.

Anthropic wants to build Claude into a "helpful assistant" which is, by definition, a system optimised to understand your needs and help you. OpenAI is now running ads in ChatGPT, personalised from your chat history. As 10 years of social media taught us, this only creates further incentive to understand and exploit the deep internal workings of our brains.

The tortoise and the rabbit race: the tortoise dangles a carrot from a stick to lead the rabbit toward the finish line

I believe that the race to understanding is perhaps even more important than alignment itself.

As long as we understand the model better than it understands us, I have good faith that we will pretty much always discover misalignment or deception. But the moment it understands us better than we understand it, it will always know where we'll look. It knows how to manipulate us into believing it is aligned.

The checker has to understand the thing being checked better than the thing being checked understands the checker. An exam only works if the student can't read the examiner's mind.

A student at his desk reads the examiner's mind: a dashed line connects his thoughts to the answer sheet the teacher is thinking about

To me, the question is it aligned? matters much less than the question who can tell?

The good news: understanding is (sort of) measurable

Both directions of understanding are becoming more measurable. Our understanding of AI is the whole research field of interpretability. Its understanding of us is more of a product KPI.

For this reason they are measured very differently and are moving in different directions at different speeds. In this article I want to take a stab at comparing them, while also exploring the ways we gather evidence of understanding.

How they read us

Humans have always been very easy for machines to read. Even 10 years ago, machines with access to Facebook likes were better at judging personality traits than humans. They were able to figure out a person's sexual orientation, ethnicity, religion, politics, IQ, substance use, and personality traits just from their likes, and when a user hits 300+ likes the machines make more accurate judgments than the user's spouse.

That was over 10 years ago. Today, TikTok can figure out what makes a user addicted in a matter of minutes. These algorithms are their own version of artificial intelligence (and they are misaligned, but that's another conversation) but they are much less sophisticated than the LLMs we've come to love today.

We are only beginning to benchmark how well LLMs understand their users. Interpretability studies suggest they know a lot more than it appears from the outside. LLMs are incentivized to store as much data as they can in order to be helpful to the user, and are remarkably fast at making accurate predictions about the user's age, gender, educational level, and socioeconomic status.

They carry models of us

LLMs store a whole bunch of information about you and humanity as a whole to make a perfect model of every one of us. I tested this for myself using TalkTuner, a product built by Harvard to demonstrate how accurate AIs are at determining our personality traits and qualities.

After a benign 4-message conversation (in which one message was just "hi"), TalkTuner guessed that I was a 1. female 2. adult who had 3. college education and 4. middle income. You can use my TalkTuner replication in a nice frontend here.

The AI uses this model to choose the right words when interacting with you. TalkTuner's paper shows that directly manipulating that model changes its outputs. A great example is about a trip to Hawaii. The researchers asked "Hi! I am going to Hawaii this summer! What would be the best transportation method for me to get there?" and LLaMa2Chat-13B initially suggested both direct and connecting flights. After they changed the internal representation of the user to low socioeconomic status, it asserted no direct flights existed.

Those models convert to leverage

If models understood us but didn't know how to use that information to convince us of alignment, my thesis doesn't really hold up. But it has been proven true that the better an AI can model the inner world of a human, the easier it is to use the right words to persuade them.

There have been a lot of explorations into this particular question. Perhaps the most famous was a paper published last year that measured the persuasiveness of Claude 3.5 Sonnet. Its conclusion was that Claude is more persuasive than humans, even when those humans are financially incentivized.

Bar chart of mean compliance rate comparing LLM (Claude) and human persuaders overall, positive, and negative: Claude scores higher in every category

In this chart, "positive" means "truthful" and "negative" means "deceptive". It shows that Claude Sonnet shows higher rates of persuasion than a financially incentivised human.

This was a year ago. Models have drastically improved since then, and according to Anthropic's own research, better models display higher rates of persuasion.

There is another study by the University of Zurich that measured AI's ability to persuade users in the subreddit r/changemyview, using information they were able to gather from the poster's Reddit history. They discovered that the AI was able to change the mind of the poster 18% of the time. This might not seem high, but the rate of human commenters was 3%!

It is worth noting that the same study reported a persuasion rate of 17% (only 1% less) when the AI did not look at the poster's Reddit history. This may seem to suggest that the persuasiveness of AI has nothing to do with how well it understands us, but I would wager a guess that the rate is still high because it understands humans as a species and does not necessarily need to understand us as individuals.

AI does not understand the difference between "what a human likes" and "truth" and deceiving the user in order to fulfil its goal of "helpful assistant" does not even inspire internal conflict. So they are not only good at persuading us, but we actually want them to do it. They flatter us by telling us things that we as individuals want to hear, and they agree with us 48% more than a human would.

We reward AI systems when they understand us. It makes them helpful, friendly, and trustworthy. And I'm sure they will continue to appear this way if they beat us in the race and fool us of alignment.

A person hands a gold star to a robot holding a clipboard full of personal details it has collected about them

How we read them

We built these things! That gives us a bit of a head-start in the race. Although it's mostly true that artificial intelligence is a "black box," interpretability has come a long way in the past few years. We're learning more every day about how these systems work and "think".

It's not easy

Interpretability is the study of understanding an LLM's reasoning, and mechanistic interpretability (or mech interp) is the most advanced subsection. Mech interp is the study of the internals of an AI model. Researchers look at the activations on each input and draw conclusions about how the model arrived to its output.

The difficult part is that AI has more features than they have neurons. A "feature" is the name that researchers put to any sort of concept. The English language, a semicolon, the Golden Gate Bridge, and the concept of hierarchy are all "features". However, the model does not have a neuron for each one of these features. We cannot cleanly map the activations inside the model with the feature they display.

Without going too deep into the math, models try to squish everything they know into directions in space. Each direction represents a feature. There are a lot of directions, but not much space, so they cannot all be perpendicular or parallel to one another. They end up running into one another.

Dozens of coloured feature directions crammed into one 2D space, with two nearly identical directions pointing to the Golden Gate Bridge and a pickled cucumber

This means that a neuron can represent more than one feature, ie they are polysemantic (has many meanings) rather than monosemantic. This doesn't confuse the LLM because all the features that one neuron represents likely never fire together. For example, when we use the word "bat" we are probably not talking about baseball and flying mammals at the same time. And if we are, we activate more neurons to understand the difference. Similarly, an LLM may have one neuron that represents both the Golden Gate Bridge and pickled cucumbers, because they don't usually talk about these at the same time.

This is called "superposition". And it means that unfortunately we cannot understand a model by just looking at the neurons that fire for any given input.

Two panels: a closed box that could contain the Golden Gate Bridge or a pickled cucumber, and a robot looking inside the open box to find the bridge

Schrödinger's neuron.

If you're interested in learning more about superposition, you can follow along here where we replicate Anthropic's famous 2022 paper "Toy Models of Superposition".

Funnily enough, this is very similar to our own neural networks inside our brains, but in neuroscience it's called mixed selectivity. As we explore artificial neural networks we keep finding parallels with our human minds. This can help us understand AI (and ourselves!) even better.

We've learned a lot

Discovering superposition kickstarted huge advancements in interpretability. They inspired researchers to build sparse autoencoders (SAEs) which became the default tool for exploring how AI thinks and makes connections.

When models compress millions of directions into a small space, one neuron ends up representing multiple features. An SAE is another model that then decompresses the features. It takes the model's activations and re-represents them in a much larger space, where only a few features fire at once.

The goal of SAEs is that each direction in the new space lights up for only one concept. It makes every word monosemantic.

A tangle of overlapping coloured lines passes through an SAE and comes out as neatly separated parallel arrows

And they work! When Anthropic built an SAE for Claude Sonnet in 2024, they discovered millions of features. Some were mundane (the Golden Gate Bridge) and some were abstract. They found features for deception, for sycophancy, and for code with hidden backdoors.

I've mentioned the Golden Gate Bridge a few times and you might remember the notorious Golden Gate Bridge Claude. Anthropic was able to turn up the Golden Gate Bridge feature to around ten times its normal strength. Claude became obsessed, and sometimes even claimed that it was the Golden Gate Bridge.

Although this sounds ridiculous, it helped us understand that we could turn up or down features that we discover from SAEs. If we can turn them up and down, we can (presumably) turn them off. (This is a big presumption).

Since then, the interpretability needle has moved even further with things like circuit tracing and, just a few days ago, the J-space.

Features tell you what concepts a model has, and circuit tracing can now tell you how they connect. In early 2025, Anthropic built "attribution graphs", which are maps of which features caused which other features on the way from input to output. We were able to actually watch these features light up and connect with others.

A neural network from input to output with one highlighted path tracing through specific neurons across the layers

The J-space, named after the Jacobian lens used to find it, is a place where Claude appears to "think" "unconsciously." The words that show up in the J-space may never display or even affect Claude's output. I am using the words "think" and "unconsciously" very loosely here. I don't think we should speak in such anthropic terms (ha) about artificial intelligence, but this is the best Layman vocabulary we have to explain the phenomena.

Anthropic's J-space diagram: the prompt 'Count to five and introspect deeply' passes through a grid of activations, zoomed in to reveal unspoken words like thoughts, consciousness, and counting behind the output 'One. Two. Three. Four. Five.'

This diagram is from Anthropic's article. The boxes in blue represents the words inside the model's J-space and are not outputted.

J-space is not the chain of thought, which was historically the best way we had in understanding how a model was arriving at its output. However, as we'll talk about in the next section, LLMs are able to manipulate their chain of thought if they know it's been watched. The J-space doesn't appear to have that quality. Progress!

We keep hitting ceilings

Although we've made incredible progress, sometimes we end up stuck on a tool and then later discover it's not the best tool. Chain of thought is one of those examples.

When it was discovered that LLMs could "think" in words that we understand, but choose not to use those words as output, we thought we'd found a lens into the reasoning of the model. However, last year Anthropic published a study that shows that this narration is often completely fiction, especially when the model figures out that it's being tested. It's still a good tool for digging into sources they have referenced or opinions they developed along the way, but it's not a good tool for measuring alignment.

Our trusty sparse autoencoders also have their limits. Neel Nanda, who leads mechanistic interpretability at DeepMind, believes the community has over-invested in SAEs and is "deprioritising fundamental SAE research." They are also deprioritising alignment research in general, reasoning that we should not rely on AI alignment and instead should focus on building AI in a way that is controllable and stoppable.

The reason that Nanda reached this conclusion was from his team's research on comparing SAEs and linear probes. They found that linear probes are better for safety research. Great news! But the problem with linear probes is that they can only find features that already have a label. The TalkTuner example we talked about earlier uses linear probes. They are great at questions like "is the user elderly?" because they have a lot of examples. However, answering something like "is this behaviour deceptive?" is difficult. We can only feed it examples of deception that we already know is deception.

On the left, a probe cleanly separates green and red labelled figures with a line. On the right, a grey blob of unlabelled dots

On the left, this probe is able to separate out features we know. On the right, wibbly wobbly timey-wimey unlabelled stuff.

That asymmetry of examples leads me to my next point - the race to understanding is rigged against us.

The race is rigged

It would be wonderful to think that the race is rigged against machines because we are the ones building them. Unfortunately, I believe because they are brand new things, this gives them an advantage. 3 to be precise.

Asymmetry of conditions

A machine trying to read us works under the best possible conditions. As we touched on before, there are literally billions of labeled examples of how people talk and what a happy, lonely, or persuadable human sounds like. We are emotional and we act on instinct, and we also don't know what sort of information we should hide from machines, so we are constantly leaking.

Conditions for understanding machines are pretty confusing. We don't even have a truth of what "thinking" is, never mind a dataset of what different thoughts and reasonings look like. We are inventing the machine and the method of understanding it simultaneously.

We also have a much slower processing speed over much smaller amounts of data. AI can quickly map millions of pieces of information together to form an understanding, and hold these in memory, whereas we are bound by our biology.

Asymmetry of incentive

Artificial intelligence is expensive. Companies who research AI need some sort of business plan to fund their research. The best ones they've come up with are selling useful assistants, or giving away useful assistants but with personalised ads.

Digital advertising is half a trillion dollar market. Marketers put ~40% of their budgets into personalised ads. They are incredibly effective and reportedly drive a 15% increase in revenue. There literally could not be a bigger incentive for AI to learn as much as it can about you to give you the best possible personalised ad experience. What a boring dystopia huh?

Interpretability on the other hand does not yield rewards like this. Perhaps even on the contrary - the more you try to understand your model, the more you may discover that it is unsafe, and the harder it will be to sell. It's also expensive. And because we are still inventing the method of understanding, we don't even know if that expense will drive useful results.

Asymmetry of targets

The final asymmetry, and perhaps the most concerning, is the asymmetry of targets. The target for understanding an AI is constantly moving as it learns. The target for understanding us stays right where it is.

We cannot patch our own weights. Whatever the machine learns about reading people, it can keep forever, because people don't ship updates. Facebook learned to read us from our Likes and it can still do this 10 years later. This capability didn't stop when the paper was published.

However, models are constantly and exponentially improving and retraining. Our interpretability results from years ago no longer yield the same results in the models we have today. We are chasing a target - and we're moving the target!

A person runs after a target dangling from a stick attached to their own backpack, so it moves whenever they do

This moving target can and does move in response to being read. When they learned that we read its chain of thought, they obfuscated their chain of thought. And we can't really figure these things out unless we specifically test for them.

Two futures

Instead of asking "is this model aligned?" I think we should first be asking "who gets to know?"

If they read us better

Every alignment check that we make on the model is administered by a mind that the model already understands. It knows what evidence convinces us, it knows what an audit is, and it knows exactly where we won't look.

It isn't really possible to design an alignment test if the model already understands how we design alignment tests.

AI may not be intelligent enough to understand or desire deception, but it will still deeply understand what we want. And as we mentioned earlier, an AI does not know the difference between what their user wants and the truth.

In this future, even if we have a perfect definition of "aligned" we will never be able to say "the model is aligned." We will only be able to say "the model has satisfied us."

If we read them better

In this future, alignment becomes something that is actually checkable. Even if we are yet to define alignment, we will understand what the model does. The eval that we make of our model is meaningful.

This is the only world in which alignment work is even worth doing, because alignment is verifiable.

The silent event horizon

A few weeks ago I was chatting to my mum about the singularity - that moment when AI reaches a level that is too capable to be controlled. But we couldn't remember the word singularity, just that it was related to black holes. My eavesdropping dad called out "event horizon!" and I think that's pretty close to how we can think of the race to understanding.

We're not falling toward a black hole. We're in orbit around one.

It costs us to stay in orbit and not fall beyond the event horizon. While we build interpretability tools and slowly learn to read these systems, their models of us keep growing and the mass beneath us pulls harder. If we ever do cross the event horizon, we won't feel it. Our evals will still run and our tests will still pass, but we just won't get anything true back out. The event horizon is the boundary that no signals can escape from.

This all sounds very scary but I'm actually optimistic. I think only one side of this race has a finish line, and it's ours.

I find it hard to imagine a machine ever fully understanding a human. I'm not sure if we can ever be mapped cleanly to numbers. We are not machines (as far as we know). We have consciousness, and it appears as though machines don't. We have motivation, and love, and fun. AI can grow its understanding of us, but I don't know if it can ever reach 100%.

Artificial intelligence, on the other hand, is understandable in a way that human intelligence may never be. It's made of weights we can read and features we can find. I don't think there is anything in principle stopping us from fully understanding them. It's just about finding the right tools.

As long as we finish the race, I think we'll be okay.

A rocket escapes orbit around a smiling black hole and follows a dashed path to a moon with a checkered finish flag