Publishers and e-commerce brands have begun testing LLM-honeypotting — a content protection technique against crawlers from OpenAI, Google, and Meta. The method is built on a classic information security approach: instead of simply blocking bots, they lure them into traps that increase the cost of scraping and pollute training models with useless data. The goal is to change the economics of large-scale content collection and make it unprofitable.
How LLM-honeypotting works
LLM-honeypotting is an adaptation of classical deception methods from cybersecurity to protect against AI crawlers. Simon Whistow, co-founder of CDN provider Fastly, explains the logic: if an attack on a system costs more than the potential gain, the attacker's business model stops working. Applied to language models, this means turning certain website visitors — obvious bots or unwanted crawlers — into targets for active defense.
There are several implementation options. The first is to force bots to perform additional work through proof-of-work tasks or imperceptible download slowdowns: large botnets face real computational bills on every page, while live users barely notice the changes. The second option is a trap in the form of an infinite content maze: the system generates an infinite number of plausible but meaningless pages that only bots see, wasting their time and computational budget.
The third approach is model and search system poisoning: bots are fed statistically coherent but factually meaningless content. When such data enters the training dataset, it degrades the quality of language model responses or triggers hallucinations, undermining trust in systems that never paid for content access.
Who is testing AI scraper protection technology
Currently, LLM-honeypotting is being tested by a small group of publishers and e-commerce brands, not the mass market. Whistow does not name specific clients, but notes that some major e-commerce players are already using this methodology with initial positive results. Interest comes from both traditional news publishers and e-commerce platforms.
The main goal is not theatrical punishment of big AI companies, but changing the economics for the long tail of scrapers who now collect data at virtually no marginal cost
Rather than simple blocking, the goal is to make every request cost scrapers real money while simultaneously reducing the quality of the data they obtain. If they can spend, say, ten million dollars in startup funding for a scraping service in a single crawl, the business model of such projects becomes unviable, and the market shrinks.
Criticism and limitations of the method
Not all experts consider LLM-honeypotting an effective solution. Frederick Yan, co-founder of Centennial and an AI solutions developer, believes the technique is too simple to detect and bypass. In practice, maze-like pages or meaningless content often aren't even shown to hidden crawlers because the defense fails to recognize them as bots on the first visit.
In his view, this is more of a marketing move than a real path to the goal. The only effective way to change publishers' position is to create friction at the protection level itself, not after a bot has already entered the site. Additionally, infinite content mazes are not free: their creation and maintenance cost more than simply blocking a bot.
Whistow agrees that this is not a universal solution and not suitable for every publisher. If a site is simple and cheap to maintain, the math doesn't work. For large, complex sites with expensive page generation and real revenue, the logic is different. A combination of increased costs for scrapers and the psychological benefit of "resistance" can justify the experiment — especially if the publisher already uses an edge platform like Fastly, Cloudflare, or Akamai, where additional computation costs less.
What this means for brands and their influencer marketing strategy
Publishers' fight against AI crawlers is changing the rules of the content market. Brands building promotion strategies through blogger advertising and media buying will face a new reality: some platforms will become harder to access for automated audience reach and engagement analysis. This increases the value of manual blogger selection and quality performance analytics for integrations — tasks where agencies with expertise and proprietary databases retain an advantage over automated tools. For example, the ETC team uses proven methodologies for audience assessment and KPI forecasting that rely not only on open data but also on direct contacts with platforms and campaign history.
For brands, this is a signal: investments in partnerships with influencer agencies that have proprietary data and relationships with bloggers become more reliable than betting on aggregator platforms dependent on mass scraping. In conditions of active content protection, access to quality media plans and proven integrations becomes a competitive advantage.
Frequently asked questions
What is LLM-honeypotting in simple terms
LLM-honeypotting is a content protection technique where websites lure bots from OpenAI, Google, and other AI companies into traps: infinite pages with meaningless content or tasks requiring expensive computation. The goal is to make data collection economically unprofitable for scrapers.
Which brands use AI crawler protection
Specific names are not disclosed, but the technology is being tested by major e-commerce brands and news publishers, primarily those working with expensive page generation and edge infrastructure. There is no mass adoption yet — this is an experimental practice for players ahead of the market in content protection.
How effective is LLM-honeypotting against scraping
Effectiveness is disputed: supporters argue that the method changes scraping economics and can spend millions of dollars of botnets' startup capital in a single crawl. Critics consider the technique easily detectable and bypassable, more of a marketing move than real protection — especially if the system fails to recognize hidden crawlers at entry.
In brief
- Publishers and e-commerce brands are testing LLM-honeypotting — traps for AI crawlers that increase scraping costs and pollute training models.
- The technology uses three approaches: proof-of-work tasks, infinite content mazes, and data poisoning with statistically coherent but meaningless content.
- A trapped bot can view 20 million pages over five days, spending scraping service startup capital of tens of millions of dollars in a single crawl.
- The method is being tested by major players, primarily on Fastly, Cloudflare, and Akamai platforms, where additional computation is cheaper.
- Critics consider LLM-honeypotting easily detectable and bypassable, more of a marketing tactic than effective protection.
- For brands, publishers' strengthened content protection increases the value of influencer agencies with proprietary databases and direct platform relationships.
* Instagram and Facebook are owned by Meta, recognized as an extremist organization whose activities are prohibited in Russia.
Want to see where the market is heading before your competitors do? The ETC team builds a media strategy and media plan for your niche — with reach forecasts and KPIs fixed in the contract.