Generative AI promises a revolution. It offers the ability to synthesise everything that had ever been published, at speed, at scale, and pretty close to free (subsidised for now by the massive capital inflows and competition for compute share). For a time, it has felt like magic. Anyone with a prompt and a bit of imagination could access the world's accumulated knowledge, translated into copy, strategy, headlines, or code.
But beneath that initial thrill was a more uncomfortable reality: these models had already consumed everything they could legally (and sometimes questionably) reach. The public web. Wikipedia. Blog archives. Forum posts. Code libraries. News articles. Instructional how-tos. YouTube transcripts. Everything the open internet had made visible.
And now, we are seeing the consequences of that saturation.
When every AI engine has trained on the same corpus of open content, we stop being surprised by what they return. The outputs flatten. The insight blurs. Competitive advantage starts to erode. Everyone is drawing from the same well.
We've reached what you might call the visibility ceiling. Too address this, and create advantage, the LLMs are pursuing two options. One is to generate synthetic data to learn from (but that is like repeatedly photocopying a copy - reductive and blurry), the other is to unlock what remains hidden.
In this next phase, differentiation won't come from better prompting or smarter prompt engineering. It won't come from downloading the latest Chrome extension or ChatGPT plugin. It will come from something far more foundational: access to data that the models haven't seen yet.
The hidden. The proprietary. The gated. The private.
What I call data from the edges.
And right now, we're watching those edges become visible, for a price.
The Internet is Changing Shape
In the last 60 days, two high-profile moves have quietly rewritten the future of discoverability online, though most marketers haven't grasped the magnitude of what just happened. They have been surprisingly under-reported.
First, Instagram. For over a decade, it's been a beautifully constructed walled garden. One of the most visually rich archives of culture, consumption, lifestyle, and commerce has been entirely closed off to Google. But that changed in July. Meta removed the 'noindex' directive from professional accounts, allowing Google to begin crawling public posts: Reels, carousels, and even captions. Just like that, one of the richest troves of unstructured, visual-first consumer content became searchable.
Second, Reddit. After licensing its data to Google in a landmark deal reportedly worth $60 million per year, Reddit then sued Perplexity AI for allegedly scraping its platform in breach of its robots.txt restrictions. The message was unmistakable: if you want our data, you'll have to pay for it, and use it on our terms.
At first glance, these two stories may seem unrelated. One is a strategic platform partnership. The other is a legal challenge to an AI startup. But they point to the same underlying shift: the battle over the next generation of indexable, monetisable, and strategically useful data has begun.
And the most valuable data isn't what's already out there. It's what's just out of reach.
The Model is Full. Now What?
It's worth pausing to understand just how complete the large language model (LLM) training sets already are. By mid-2024, nearly all major LLMs had incorporated the open internet up to that point. Some trained on Common Crawl, others on academic and code repositories, others on social forum threads. A few snuck in entire media libraries until lawsuits stopped them (many more of these to come...). What's left?
What's left is what they couldn't get to. What sat behind paywalls, inside apps, or on platforms that explicitly blocked indexing. What lived in private channels, niche forums, and brand-owned environments. What was too unstructured, too ephemeral, or too human for traditional crawlers to parse.
And that's the data that's now coming into focus.
The move to index Instagram wasn't about giving users better access to pretty photos. It was about giving Google (and by extension, its LLM Gemini) access to a corpus of visual and cultural data it never had. A hundred thousand products being used, worn, cooked with, reviewed, in motion, in situ, in context. All previously invisible to search. Now, slowly, made visible.
The Reddit lawsuit, conversely, isn't about stopping bad actors. It's about asserting ownership over data that's suddenly seen as economically valuable. If Google will pay for the privilege, then so must everyone else.
We've entered the Index Economy.
What This Means for Marketers
For most of the last two decades, digital marketing strategy has been built on an assumption: make yourself findable. SEO. SEM. Social content. Optimised journeys. We built pathways for consumers to find us, engage with us, click through, subscribe, convert. But now, the rules have changed.
When 70% of global searches end without a click (as Similarweb reports) and AI Overviews push that number to 83%, we no longer own the post-click experience. Users get answers inside the search engine. Content is re-aggregated and re-summarised before it ever reaches your site. And the AI doesn't always cite the source. Even when it does, users often don't click.
The implication is that the content you create is now part of a larger training corpus. It powers someone else's model. And if your data isn't distinct, you simply disappear into the average. So how do you stay distinct?
You do it by owning what the models can't replicate. By producing and protecting the information they haven't seen yet. By activating your brand through what they can't generate.
You do it with edge data.
That includes your owned customer journeys. Your gated insights. Your influencer collaborations. Your proprietary surveys, community forums, loyalty programs, and live event content. It includes social media posts that previously never left the feed, and are now being indexed for the first time.
Your edge is not just in the content you publish, it's in the context you control.
Edge Data Is Strategic Now
For a long time, edge data was considered secondary. Customer support logs. In-app feedback. Loyalty engagement patterns. CRM sentiment notes. Not sexy. Not scalable. Not public.
But now, that invisibility is a strength. In a world where AI tools generate from what's already seen, anything unseen becomes a signal of differentiation.
Platforms understand this. That's why they're building walls around their own edge data and inviting others in for a fee.
Reddit isn't just licensing access to its threads. It's licensing the social consensus, the back-and-forth debate that trains sentiment models. Instagram isn't just being indexed, it's becoming a structured archive of everyday product demonstrations and cultural expression.
The next frontier of brand strategy will depend not just on what you publish, but what you uniquely possess.
And more importantly, on what you're willing to share, and with whom.
The Age of the Platform as Publisher
The other, quieter shift is how platforms are repositioning themselves. No longer simply distribution pipes, they're now becoming publishers of context.
Google doesn't just index content, it curates it. Meta doesn't just host creators, it owns the underlying engagement infrastructure. Amazon doesn't just run retail ads, it's training models on shopping behaviours.
And in each case, they're beginning to license their advantage.
Instagram becomes a dataset. Reddit becomes a data vendor. TikTok quietly pilots API access to trending tags. LinkedIn begins surfacing video captions to third-party search.
This is the strategic unbundling of the social web. The new game is no longer attention. It's access.
The Opportunity (and Responsibility) for Brands
The good news is this: brands are not powerless. In fact, many are better positioned than they realise.
Your owned first-party data about your customers, your products and your usage trends is untouched by any LLM. Your CRM, your customer feedback, your transactional insights: none of that was in the training sets. Your influencer partnerships, when structured for discoverability, can now play a dual role: in-feed engagement and search visibility. Your product how-tos, community interactions, loyalty journeys, and even customer service logs hold more insight about your brand than any search engine summary could dream of producing.
But it must be captured. Protected. And activated on your terms.This isn't just a call for data hygiene. It's a shift in strategic mindset.
In a world where everything visible has already been absorbed, only the invisible remains valuable.
And as platforms race to turn the previously unindexable into structured, monetisable data streams, marketers must decide: will we be passive participants or active players?
If you take one thing away, let it be this: we've moved beyond visibility. The future is about ownership.
Ownership of data. Ownership of context. Ownership of how your brand shows up in a world mediated by machines.
You won't win by shouting louder. You'll win by knowing more—about your customer, about your product, and about the edges no one else is watching.
Because when the models have read it all, the only source of advantage left is the part they haven't seen yet.
That's your edge.
Brandflow is written by Justin Billingsley, who has spent his career on all three sides of the industry's table: senior client, global agency leader, technology founder. First published 14 July 2025 in the Brandflow newsletter on LinkedIn.

