All journal articles

News

Perplexity’s new retrieval benchmark makes one AI-search layer easier to inspect

Q2D-Web evaluates which documents enter an agentic search system’s candidate set. It is not a citation report—but it sharpens the distinction.

September 11, 2026 5 min read

DIGRAPH / JOURNAL
A physical archive of paper source records arranged around one cobalt-blue reference marker
SIGNAL STUDY 06 · Before an answer can cite a source, a retrieval system must decide which documents are candidates.

AI search is becoming an operating surface for brand teams. The practical question is not whether a single model response changed—it is whether the change reveals something your team can verify and improve.

What Perplexity released

On September 9, Perplexity Research introduced Q2D-Web, a benchmark and public leaderboard for first-stage retrieval in agentic retrieval-augmented generation systems. The accompanying paper describes 70,000 agentic queries, a 190 million-document web corpus, and three sets of relevance judgments. In plain terms, it evaluates an early but consequential step: which documents a system puts into the pool before later ranking, synthesis, and answer generation. Perplexity describes Q2D-Web as a research benchmark, not as a change to the sources or citations shown in its consumer product. That boundary is important.

Why the first retrieval stage matters

An answer system cannot read every page on the web for every question. A first-stage retriever narrows the field to a large candidate set; later components can rerank, select, cite, summarize, or omit those candidates. If a useful page is absent at that early stage, it may never be available to the later steps. If it is present, that still does not guarantee a citation, a mention, or a favorable description. Q2D-Web therefore makes a hidden selection problem more measurable without making it visible in any individual answer.

The benchmark is larger than a simple link test

The paper says its corpus contains 190 million web documents and its queries were reformulated from consented production-system queries after PII filtering. It also uses deeper relevance labels than a benchmark that marks only one known good document. Those design choices are intended to test whether a retriever can find many genuinely relevant documents, including ones that look plausible but do not answer the query. The published leaderboard can help retrieval researchers compare approaches under a shared setup. It does not show which domains Perplexity currently ranks, cites, or sends traffic to in a live user session.

Do not confuse retrieval, citations, and recommendations

For brand and content teams, these terms should remain separate. Retrieval is an internal candidate-selection step. A displayed citation is a user-facing link attached to a response. A recommendation is the answer’s judgment about fit or preference. One page can be retrieved and never shown; cited without its brand being recommended; or used alongside sources a reader never sees. Treating a benchmark score as proof that a page will earn more citations would be the same category error as treating an answer mention as proof of a ranking rule.

What this changes for AI-search research

The release gives researchers a more concrete way to study the web-retrieval layer beneath agentic answers, and the independent coverage has correctly framed it as a benchmark rather than a new search interface. That matters because source behavior is often discussed with too little precision. Q2D-Web may improve how retrieval methods are evaluated over time; it does not establish that any platform has changed its current source selection. Teams should watch for replicated findings, model results, and product documentation before translating benchmark movement into a visibility forecast.

The practical response is an evidence record

Continue measuring the signals buyers can actually encounter: the exact prompt, market and date, full answer, sources displayed, brands named, recommendation language, and relevant referral behavior. Keep an adjacent record of the public pages and independent sources that substantiate important claims. If an answer changes, investigate whether the user-facing citation, answer framing, or supporting evidence changed before assuming an unseen retrieval layer is the cause. This produces a decision-ready record while acknowledging that most systems do not expose their complete retrieval pipeline.

What remains unknown

Q2D-Web does not reveal Perplexity’s production ranking logic, its final citation policy, the proportion of users or queries represented by the benchmark, or the effect of a higher benchmark score on user outcomes. It also cannot tell a publisher how to be retrieved for a particular query. The useful lesson is more restrained: strong source visibility has several stages, and a public benchmark can illuminate one of them. The right response is clearer measurement vocabulary—not a new checklist or a promise that an optimized page will be selected.

Bottom line

Perplexity has published a substantial new benchmark for evaluating the document-retrieval stage of agentic search. It is relevant to anyone studying how web pages become available to answer systems, but it is not evidence of a live citation or recommendation change. Brand, SEO, and content teams should keep their attention on observed answers and sources, while using the release as a reminder to distinguish candidate retrieval from what a buyer ultimately sees.

Sources