Skip to content
astra.buzz
Go back

Perplexity is citing content farms as sources

· 5 min read · 943 words

Perplexity describes its product as an answer engine that researches the open web and returns accurate answers “backed by citations.” It says those citations make its answers checkable rather than opaque. That promise is the whole pitch. If the retrieval system cannot tell research from mass-produced marketing copy, the citations become decoration.

A new report from Trellner tested that promise with 380 buyer-intent software categories. The researchers sent every category to Perplexity’s sonar and sonar-pro models through OpenRouter, then kept every URL the models retrieved. Across 760 calls, they recorded 7,534 citations from 2,055 domains. According to the published dataset, 59.8% of the citations pointed to domains ranked below 100,000 on the Tranco top-million list. Nearly a quarter pointed to domains outside that million entirely.

Tranco measures popularity, built by combining several web rankings over 30 days. Truthfulness requires separate evidence, and Trellner says so in its limitations. The ugly part is what the researchers found when they opened some of those sources and looked at the business models behind them.

A marketing blog outranked Gartner

Guideflow was Perplexity’s third most-cited domain in the test, with 194 citations across 96 of the 380 product categories. Gartner had 158. Guideflow sells interactive product demos and describes its own blog as a place for demo automation advice and company news. Yet that blog currently stretches to 181 index pages, packed with posts such as “project profitability software,” “critical path software,” and “visual work instruction software.”

Trellner found a different Guideflow URL cited for each of those 96 categories. The researchers also reported that Guideflow competed in none of the categories they tested. Publishing thousands of search-oriented listicles is ordinary content marketing, and Guideflow makes no deceptive claim by doing it. Perplexity is responsible for deciding that a demo vendor’s marketing blog belongs ahead of Gartner in the evidence base for software recommendations.

The failure belongs to the system selling the answer as researched and checkable. A retrieval engine should understand that a vendor’s content funnel and an independent product analysis carry different conflicts and deserve different weight. Perplexity’s stack treated topical coverage as authority.

Three sites produced 215,128 buying guides

The stranger evidence sits farther down the citation table. WifiTalents, Worldmetrics, and Gitnux received 181 citations across 41 categories. Trellner’s sitemap analysis counted 215,128 URLs under the three sites’ /best/ paths. The sites published 215,128 guides; Perplexity cited their domains 181 times.

The three sites look coordinated, although Trellner correctly stops short of claiming proven common ownership. The report found the same Cloudflare nameservers, nearly identical templates and navigation, and tiny blogs that promote the other brands in the group. Their current homepages make the scale visible without any inference. Gitnux says it has 72,001 best lists, while WifiTalents says it has 72,729. Worldmetrics uses the same structure and advertises software advice alongside paid market research.

Worldmetrics and Gitnux also give their homepages the HTML title “Facts & Grounding Page.” Their matching meta descriptions call each site an independent research company with verified facts in a “machine-readable record.” Grounding is the retrieval step that fetches outside material for a model. Private intent remains unknown. The wording is plainly written for software systems as well as human buyers.

The production defects are just as revealing. Trellner saved the three sites’ versions of a page about project estimation software. Gitnux ranked Saviom first. Worldmetrics ranked Float first. WifiTalents also chose Float, but the rest of its top five differed from Worldmetrics. The pages credited nine people across the three brands and advertised expert review. All three displayed broken template text promising another update “within the next” 26 or 40 days.

Nobody has to prove those rankings wrong to see the source-quality failure. Perplexity cited pages produced by a network whose scale, duplicated structure, machine-facing language, conflicting verdicts, and visible template bugs should all have triggered scrutiny.

A correct answer can still carry rotten evidence

The recommendations may be perfectly reasonable. Trellner left the causal question untested: whether removing these sources would change Perplexity’s answers. The study covered one search stack on one day, and its 380 categories were chosen by the researchers instead of sampled from real user traffic. Sonar and sonar-pro shared almost all of their citation lists, which points to a common retrieval system rather than independent confirmation.

Those caveats narrow the finding to a snapshot of one engine. Even within that boundary, Perplexity’s evidence chain failed a basic provenance test. An answer can name a good product for bad reasons. A citation identifies the source. Trust still depends on the quality and conflicts of that source. Perplexity markets citations as the feature that turns a chatbot response into checkable research, which makes provenance part of the product and deserving of product-level quality control.

This is where AI search can become worse than the web search it claims to replace. A list of links forces the user to see the domains and choose what to open. A conversational answer puts the conclusion first and reduces source review to an optional click. The interface transfers judgment from the reader to the retrieval system. If that system rewards page volume, repeated category phrases, and machine-readable formatting, publishers will manufacture exactly those things.

The web already ran this experiment. Search engines rewarded keyword coverage, so content farms buried useful pages under material written for ranking systems. AI search now offers the same economic prize with a smoother disguise. The spam can disappear from view while its claims survive in the answer.

Perplexity needs a retrieval layer that can recognize who published a page, why it exists, what conflicts shape it, and whether another credible source agrees. Until it can do that, “backed by citations” means the machine found some URLs. Verification requires far more than that.


Share this post on:

Previous Post
Trump's Justice Department left an ICE shooting out of its federal case
Next Post
CMS is holding Texas hospital care hostage for $27 million a day