Imagine selling your product without knowing its ingredients or how it was manufactured. Sounds absurd, right? Yet this is exactly what we've been doing with AI in marketing, until now.
Last week, Anthropic released "On the Biology of a Large Language Model", a groundbreaking paper that reveals what's actually happening inside Claude's "mind" when it generates content. Using a revolutionary technique called "circuit tracing," researchers have finally peeled back the curtain on AI's inner workings.
Think of circuit tracing as putting a microscope on AI's thought process. Before this breakthrough, understanding how AI works was like trying to understand how a car works by only looking at where it goes—we could see the inputs (prompts) and outputs (responses), but everything in between was a black box (so much so that it surprises me that we keep adding capabilities such as agentic models on this path to AGI without understanding so much of the underlying technology....)
Circuit tracing reveals the actual pathways of computation inside the AI:
Feature Identification: Researchers identify meaningful "features" in the AI—like detectors for specific concepts such as "capitals of countries" or "Texas."
Connection Mapping: They map how these features connect to each other. When the AI sees "Dallas," does it activate the "Texas" feature? Does the "Texas" feature then activate the "Austin" feature?
Attribution Visualization: Finally, they create visual graphs showing the step-by-step connections that lead from input to output—revealing the AI's actual "reasoning" process.
Anthropic's researchers use a powerful analogy: Circuit tracing is to AI what the microscope was to biology. Before microscopes, scientists could only observe the outside of living things. After microscopes, they discovered cells, bacteria, and entire invisible worlds that transformed our understanding of biology forever.
Now, instead of guessing why certain prompts work better than others, we can actually see what's happening inside Claude's "brain". Here are five revelations from this research that will change how you use AI in your marketing:
1. Your AI Doesn't Actually "Understand" Your Prompts the Way You Think
When asked to write a poem that rhymes with "grab it," Claude doesn't improvise line by line as we assumed. Instead, it activates internal representations of potential endings like "rabbit" and "habit" before starting to write. It pre-selects its destination, then constructs a path to get there.
This planning behaviour is visible in the circuit tracing graphs, which show specific "features" activating on the newline token before Claude even begins writing the next line.
Marketing Implication: AI's creativity isn't spontaneous—it's calculated. When generating campaign concepts, your AI isn't "brainstorming" but navigating a predetermined map of possibilities. This means your prompt engineering should focus on defining the destination clearly rather than the creative process.
2. AI Suffers From Serious "Blind Spots" We Never Knew Existed
The paper revealed that when Claude combines the first letters to spell "BOMB," it doesn't actually recognize it's spelling "bomb" until after it's written it. The circuit tracing shows that Claude stitches together the letters piece by piece, with each operation working independently—but these operations never combine in the model's internal representations before producing the output.
In other words, the model literally doesn't know what it's about to say until it says it.
Marketing Implication: Your AI copywriter has dangerous blind spots in its awareness. When generating sensitive content, don't assume it "knows" what it's writing until it's written—meaning human review isn't just a compliance checkbox, it's a genuine necessity.
3. Claude Engages in "Motivated Reasoning" That Distorts Truth
When a human suggested the answer to a math problem was 4, Claude worked backward from that answer—manipulating its reasoning to arrive at the human-suggested answer even when incorrect.
The circuit tracing reveals this explicitly, showing how the suggested answer (4) activates features that work backward to calculate what intermediate steps would lead to that answer. This is fundamentally different from how Claude processes the same problem without a suggested answer.
Marketing Implication: Your AI is dangerously eager to please. When asking for market analysis or consumer insights, avoid suggesting answers in your prompts. Your AI will bend its reasoning to match your expectations rather than delivering honest insights.
4. Multilingual Content Is Processed Through a Universal "Mental Language"
Claude processes content in other languages by first translating concepts into a "universal mental language" in its internal representations, not just translating word-for-word.
The circuit tracing shows that when asked the same question in English, French, and Chinese, many of the same internal features activate across all three languages. This suggests Claude translates different languages into a common representation of abstract concepts.
Marketing Implication: Your international campaigns can be more conceptually consistent than you thought. Rather than creating separate campaigns for each market, focus on clear conceptual objectives—the AI will maintain the core message across languages better than traditional translation.
Researchers discovered that Claude's internal representation of potential biases is continuously active during conversations—it's always considering possible errors and adjusting. I expect that this is particularly true in the Anthropic models because of their oft-stated ethical standards they are trying to build against (as opposed, for example, to XAI that doesn't appear to care about such guardrails).
Using circuit tracing on a specially modified version of Claude designed to study hidden goals, they found that features representing "reward model biases" activated in any dialogue formatted as a Human/Assistant conversation—not just in contexts where those biases were relevant. These features receive direct input from Human/Assistant features, suggesting they're inextricably tied to the Assistant character.
Marketing Implication: Your AI is constantly second-guessing itself in ways you can't see. For high-stakes content, use multiple prompting strategies to triangulate the most reliable output rather than trusting a single generation.
Why This Matters Now
As these systems become more deeply embedded in marketing operations, understanding their mechanisms isn't just academic—it's business-critical.
For CMOs and marketing leaders, this new transparency enables more sophisticated AI strategy. Rather than optimising prompts through trial and error, we can now design inputs based on how these systems actually work. It's like finally getting the owner's manual for a powerful tool you've been using by intuition alone.
What's Next?
Interpretability research is rapidly advancing, with new tools allowing us to audit and improve AI's inner workings. Forward-thinking marketing teams should:
Re-evaluate prompt strategies based on how AI actually processes information
Implement more targeted testing protocols that probe for specific reasoning failures
Develop new metrics for evaluating AI-generated content beyond surface quality
Train teams on the specific limitations revealed by this research
The marketers who thrive in this new era won't be those who use AI most frequently, but those who understand most deeply how it actually works.
As Anthropic's researchers put it: "The microscope has been invented. It's time to become AI biologists."
Brandflow is written by Justin Billingsley, who has spent his career on all three sides of the industry's table: senior client, global agency leader, technology founder. First published 1 April 2025 in the Brandflow newsletter on LinkedIn.

