Ready for a week of people unpacking the genius and failure of Super Bowl advertising?
One of the groups to opine will be System1 and then those opinions will echo between their fanboys. I write this newsletter with some caution for that very reason as there are many who support this testing method and the resulting ratings. As Campaign's indomitable Maisie McCabe wrote in 2024 about System1's "insitence that they cannot be challenged". So I'm not going to be lazy and title this newsletter "Snakeoil1" and drop opinions, as that is just unhelpful. And my goal is to be helpful to people who make decisions on creativity.
I think people support System1 because it is easy to understand the results and also because they haven't taken time to understand the methodology. As US Senator Moynihan said "Everyone is entitled to his own opinions, but not his own facts", so please allow me to describe the System1 testing process and then you decide it's veracity. The methodology may reveal for you more than System1 might intend.
The company name "System1" comes from psychologist Daniel Kahneman's framework describing how humans think. In his book Thinking, Fast and Slow, Kahneman identified two distinct modes of thinking:
System 1 is fast, automatic, and unconscious. It's how you recognise a friend's face, understand simple sentences, or feel an emotion while watching a story unfold. No effort required. No conscious monitoring. You just experience.
System 2 is slow, deliberate, and conscious. It's how you solve a math problem, compare product features, or (critically) monitor and categorise your own emotional state. It requires effort, attention, and active thinking.
The distinction matters because most advertising works through System 1. Emotional response happens automatically, without conscious processing. You laugh, feel moved, or connect with a character without thinking about why you're feeling those things.
System1 (the company) named itself after this psychological framework, positioning its methodology as measuring unconscious emotional response. But here's the methodological question: does asking people to actively monitor and categorise their emotions keep them in System 1, or does it shift them into System 2?
When marketers see System1's second-by-second emotion curves, with smooth lines tracking happiness, surprise, sadness for the duration of the spot running alongside, it would be natural to assume that the testing method is some combination of:
Automated facial coding (AI reading micro-expressions)
Biometric measurement (heart rate, galvanic skin response)
Implicit association testing (unconscious reaction measurement)
Large sample sizes (500+ respondents like traditional quantitative research)
Even the "FaceTrace" method name reinforces this impression. The visualisation implies continuous measurement. The "System1" branding suggests unconscious emotional response.
That's not how it works
System1 doesn't document who they use for recruitment, what devices they watch the creative on, how many they watch in a session or how much they are compensated. As I don't have facts for these, I'll leave the risks of this aspect of testing to one side, and focus on the method alone.
Respondents watch the ad with eight emotion options displayed continuously on screen: happiness, surprise, fear, disgust, anger, contempt, sadness, neutral. They are depicted as words and (helpfully) emojis. Respondents can either look away from the ad to find the emotion they feel while it's playing or pause the video at any point to click their current emotion (not certain which affects the results the most). After watching, they select the overall emotion they think that they felt. Sample size: approximately 150 respondents per ad.
Think about what this creates for someone who is pausing. You're watching a Super Bowl spot that is say, 30 seconds. You might pause at 5 seconds to click "happiness." Then at 15 seconds to click "surprise." Maybe once more at 22 seconds.
Another respondent in the study pauses at 8 seconds, 18 seconds, and 27 seconds. Different moments. Different emotions selected. Each respondent clicks 3-5 times during the spot.
System1 aggregates these discrete clicks across 150 people, interpolates the gaps between clicks, and generates smooth emotion curves that appear to show continuous second-by-second measurement.
The name "FaceTrace" is a hangover from when System1 used automated facial coding before they moved away from it. A NielsenIQ study ("About face: The shift away from facial coding technology." Published November 23, 2022.) found "...in what we believe is one of the largest and most comprehensive analyses ever conducted on these systems, we analysed facial expression data from over 2,000 video ad tests across 15 countries" that facial coding failed to predict advertising effectiveness with an r = 0.07 correlation between facial expressions and brain-measured emotional motivation, an r = 0.18 for positive expressions predicting sales and an r = -0.02 for negative expressions. The conclusion: facial expressions during passive ad viewing don't reliably indicate the emotions people actually experience.
But asking people to monitor and categorise their emotions creates a different problem: it changes those emotions.
When you're watching an ad while simultaneously asking yourself "what emotion am I feeling right now that I should categorise and click?" you're not in a System 1 state. You're in System 2. You're meta-analysing rather than experiencing.
This matters because emotional advertising works through uninterrupted narrative flow. The joke builds. The music swells. The story pays off. Fragmentation breaks the mechanism you're trying to measure.
Naturally System1 points to extensive validation, and I anticipate that these will soon appear in the comments below. The most substantive studies:
IPA 2009 (BrainJuicer era): Peter Field analysed campaigns where the IPA Effectiveness Databank already had business effects data. Finding: campaigns with higher emotional response achieved more "very large business effects." This established correlation but not causation and, critically, tested campaigns that had already proven effectiveness.
IPA 2021 (Orlando Wood's "Look Out"): 4,000 ads, 32 categories, £10 billion media spend. Adding Star Ratings to ESOV (media weight) data improved market share growth prediction. System1 claims accuracy "nearly doubles." But ESOV itself is a powerful predictor, the study didn't isolate whether Star Ratings alone drive incremental prediction.
The Creative Dividend (2026): System1 and Effie Worldwide published their latest validation featuring 1,265 Effie award-winning campaigns (2007-2023), the study introduces ESOC (Excess Share of Creativity) combining creative quality scores with media weight. Claims: profit growth increases "exponentially" with higher ESOC, and creativity plus media account for 60.1% of campaign business results.
The pattern is identical across all three studies.
The Effie Case Library contains campaigns that already won awards for effectiveness. The IPA Databank contains campaigns that already demonstrated business impact. Finding correlation between Star Ratings and proven winners validates that successful campaigns tend to score well. It doesn't prove the ratings can predict success before a campaign runs.
It's validation in hindsight, not prediction in foresight.
And none of these studies address the core methodological question: Do people watching ads in fragments while categorising emotions experience the creative work the same way as uninterrupted viewing?
ESOC rebrands common sense (better creative + more media = better results) as validation of a specific measurement methodology. But the measurement methodology's flaws remain unaddressed.
Understanding methodology isn't about rejecting testing, it's just about using it appropriately. Emotional response clearly drives effectiveness. The question is whether the measurement methodology preserves the emotional experience it's trying to measure.
For brand leaders: Request transparency from any testing vendor. Ask to see the actual respondent experience, not just the output visualisations. Question sample sizes and recruitment against industry standards. Understand whether the methodology measures passive viewing or active System 2 categorisation.
For agencies: Methodology critique is legitimate creative defence when it's factual. "This spot builds emotion through sustained narrative flow that fragmented viewing disrupts" is a valid observation about creative construction. Use it to protect work designed for how audiences actually watch, not how testing respondents click.
The visualisation implies precision the methodology doesn't deliver. Smooth emotion curves created from discrete clicks across different respondents at different moments aren't continuous measurement, they're statistical interpolation. Knowing this doesn't invalidate the approach, but it does change what conclusions we can draw from it.
Testing should inform creative decisions, not dictate them with some stars, lines and a number. Understanding the methodology helps us use testing as one input among many, weighted appropriately for what it actually measures.
Brandflow is written by Justin Billingsley, who has spent his career on all three sides of the industry's table: senior client, global agency leader, technology founder. First published 9 February 2026 in the Brandflow newsletter on LinkedIn.

