The scraping controversy, in specifics
The oldest and most damaging complaint against Perplexity is how it obtains the content it summarizes. In 2024, an investigation by Wired reported that Perplexity was accessing and summarizing pages from sites that had explicitly disallowed its crawler. In 2025, Cloudflare published a separate technical analysis reaching a similar conclusion: that Perplexity used undisclosed crawlers with spoofed user-agent strings — disguising its automated requests as an ordinary browser — to retrieve content from sites that block scraping.
This is the detail that matters, and it is why a common assumption is wrong. Many site owners believe a robots.txt disallow rule settles the question of whether an AI system can take their content. If the Cloudflare and Wired findings are accurate, that control was not being reliably honored. For a business, the practical takeaway is uncomfortable but important: a polite instruction in a text file is not an enforcement mechanism. If your content strategy assumes AI crawlers respect your stated boundaries, that assumption is worth testing rather than trusting.
CNN v. Perplexity: the lawsuit that raised the stakes
On 28 May 2026, CNN filed a 54-page federal complaint against Perplexity in the Southern District of New York — Cable News Network Inc. v. Perplexity AI, Inc., Case No. 1:26-cv-04427. It is the first suit of its kind brought by a television network, and it alleges that Perplexity unlawfully scraped more than 17,000 CNN works — stories, videos, and images — to power its products without authorization or payment.
The complaint makes two separate claims. The copyright claim alleges Perplexity copied and redistributed CNN's journalism to generate answers “identical or substantially similar” to the original reporting, and the filing reportedly includes side-by-side comparisons with the copied passages highlighted. The trademark claim is distinct: CNN accuses Perplexity of implying an ongoing content relationship by advertising CNN access to subscribers of its Comet Plus tier when no such deal existed.
CNN is not alone. Reporting around the case describes it as part of a broader wave, with multiple major publishers — among them the BBC, Dow Jones, and The New York Times — having raised copyright, scraping, or trademark objections against Perplexity. Because statutory damages under the Copyright Act can reach up to $150,000 per willful infringement, the financial exposure across thousands of works is not theoretical.
The dispute did not appear overnight. The public timeline reported around the filing runs roughly like this:
- Late 2024–early 2025: CNN and Perplexity enter licensing negotiations. They cannot agree on terms — reportedly including limits on how Perplexity's product accesses and uses CNN content — and talks collapse without a deal.
- December 2025: CNN sends Perplexity a cease-and-desist letter demanding it stop using CNN's content.
- May 2026: With no resolution, CNN files suit in federal court.
That sequence — failed licensing, warning, litigation — is the pattern to watch. It signals that the era of AI companies treating the open web as free training input is being contested deal-by-deal and court-by-court, and the outcome will shape what “fair use” means for everyone who publishes online.
You do not need to follow the litigation to act on its lesson. The reason CNN can pursue statutory damages at all is that its works were registered. Copyright exists automatically the moment you create something, but registration is what unlocks the strongest enforcement remedies. If your business produces original research, white papers, or proprietary written material that represents real competitive value, registration is the difference between a strongly enforceable right and a mostly symbolic one.
Citation accuracy and the hallucination problem
Perplexity's core promise is that it cites its sources, which should make it more trustworthy than a chatbot that invents freely. In practice, the citation layer is only as good as the links behind it. The underlying language models still produce plausible-sounding statements that are not quite right, and attaching a citation does not guarantee the cited page actually supports the claim. When citations point to outdated pages, paywalled articles, or content that does not say what the answer implies, the tool can lose trust faster than a traditional search result — because it presented itself as verified.
This stems from how retrieval-augmented generation works: the system pulls relevant documents at query time and composes an answer, but it does not deeply understand or independently verify the factual claims it assembles. For any business relying on AI answers for decisions, that is a review requirement, not a set-and-forget tool. It is also why being the clearly-worded, well-structured source an engine can quote accurately is a defensive advantage — the easier you are to cite correctly, the less likely you are to be misrepresented.
The CEO controversy over AI job losses
Not all of the backlash is about content. In April 2026, Perplexity CEO Aravind Srinivas drew widespread criticism for remarks on a podcast about AI-driven job displacement. He argued that “most people don’t enjoy their jobs” and framed generative AI as an enabler of individual “mini businesses,” describing the disruption — while acknowledging “temporary job displacement” — as a path to a “glorious future.”
The comments landed badly against the moment. Outplacement firm Challenger, Gray & Christmas reported more than 33,000 AI-linked tech-sector layoffs in early 2026, and the general critique was that the framing minimized the position of workers without the capital or runway to pivot into entrepreneurship. For our purposes the episode is less about the economics and more about a reputational reality: the perception of a platform is part of its brand, and it is set by its leadership as much as its product. The same is true, at a smaller scale, for every business owner whose name is attached to their company.
The trust problem inside the product
The most revealing backlash comes from Perplexity's own paying users. A recurring, heavily-upvoted complaint on the platform's community forums is that Perplexity silently reroutes requests: a user pins a specific model for a task, and the system quietly substitutes a different one without a clear notice. Users describe this as a bait-and-switch that undermines their trust in what they are actually getting. Related threads criticize aggressive changes to pricing and to what the free and paid tiers include.
Whether or not each specific complaint is fair, the pattern is the strategic point: trust is the entire product in an answer engine. Traditional search shows you ten links and lets you judge them; an answer engine asks you to accept a single synthesized response. That only works if users believe the machinery behind it is being honest with them — about its sources, its models, and its terms. Every erosion of that belief is an opening for a competitor, and a reminder that the same trust dynamic governs how customers judge your business online.
What the backlash means for your digital strategy
Step back from Perplexity specifically and the controversies resolve into four durable lessons for anyone managing their visibility.
- Decide your AI-access posture deliberately. The scraping findings show that blocking is unreliable and that the real choice is strategic, not technical. For most small businesses and personal brands, the goal is the opposite of blocking — you want to be found and cited — but that should be a decision you have made, with high-value or licensable content handled differently from marketing content.
- Protect what is genuinely valuable. If you create original, defensible work, treat copyright registration as the enforcement layer it is. The open web is now an input to commercial AI products; assume your best content will be consumed and plan accordingly.
- Be the source that is safe to quote. Because these engines misfire on citation, clarity is protection. Content that states facts plainly, is well-structured, and is corroborated elsewhere is both more likely to be cited and less likely to be cited wrongly.
- Treat AI visibility as its own channel. The disruption to traditional click-through search is real. Optimizing to be part of an answer is a different discipline from optimizing to rank a page, and the businesses that treat it as a distinct channel — measured, structured, and governed — are the ones capturing the shift instead of being stranded by it.
The backlash against Perplexity is not a reason to ignore AI search; it is a field guide to the risks that come with it. Understanding where a fast-moving platform is getting criticized tells you exactly where to be careful — and where the openings are — as you build your own presence in AI-driven discovery.
Our Research On This
Original SEMPITE studies — live queries, recorded answers, named sources. Free to cite under CC BY 4.0.
- Who Google’s AI Recommends in Sports Nutrition — 5.1% of AI citations go to brand-owned sites
- The 3 Publishers That Control Supplement AI Answers — 77% of AI supplement answers come via 3 publishers
- AI Visibility Index — 43% of Google top-3 businesses ChatGPT never mentions
SEMPITE helps small businesses and personal brands get found — in search and in AI answers.
Get in TouchFrequently Asked Questions
Is Perplexity AI being sued?
Yes. On 28 May 2026, CNN filed a federal copyright and trademark lawsuit against Perplexity in the Southern District of New York (Case No. 1:26-cv-04427), alleging it scraped more than 17,000 CNN works without authorization. Reporting describes the case as part of a broader wave of objections from major publishers including the BBC, Dow Jones, and The New York Times.
Does Perplexity AI respect robots.txt?
This is disputed. Investigations by Wired (2024) and a technical analysis by Cloudflare (2025) reported that Perplexity used undisclosed crawlers with spoofed user-agent strings to access sites that had explicitly blocked scraping. The practical implication is that a robots.txt disallow rule may not reliably prevent an AI system from taking your content, so it should be treated as a stated preference rather than a guaranteed control.
How accurate are Perplexity AI answers?
Perplexity attaches citations to its answers, which helps, but the underlying models can still generate plausible-sounding errors, and a citation does not guarantee the linked page actually supports the claim. Treat it as a research assistant whose sources you verify before acting, not a definitive authority — especially for decisions with financial, legal, or reputational stakes.
Why did Perplexity's CEO face backlash?
In April 2026, CEO Aravind Srinivas drew criticism for podcast remarks on AI job losses, saying most people dislike their jobs and framing AI-driven disruption as a path to a 'glorious future' of individual 'mini businesses.' Critics argued the framing minimized the position of workers without the capital or runway to become entrepreneurs, against a backdrop of tens of thousands of AI-linked tech layoffs reported in early 2026.
Should businesses block AI crawlers like Perplexity's?
For most small businesses and personal brands, no — the goal is to be found and cited in AI answers, which is where discovery is heading. Blocking mainly reduces visibility. The better move is a deliberate posture: welcome crawlers to your marketing content, register and handle genuinely proprietary or licensable material separately, and structure your content so it is quoted accurately rather than misrepresented.
Will AI search replace traditional SEO?
No, but it changes the job. Strong fundamentals — crawlable, authoritative, consistent content — still matter because AI engines draw on the same signals. What's new is optimizing to be cited and synthesized into an answer, which rewards direct answers, clear structure, and corroboration across the web. Treat AI visibility as a distinct channel alongside traditional SEO, not a replacement for it.
Does the Perplexity backlash mean I should avoid AI search visibility?
The opposite. The controversies are a map of the risks, not a reason to opt out. Understanding where Perplexity is being criticized — unreliable crawler behavior, citation errors, trust gaps — tells you precisely where to be careful and where the openings are as you build a presence in AI-driven discovery.
Leave a Comment
Have a question or something to add? Drop a comment below.