Allow PerplexityBot in robots.txt so Perplexity can index you, and verify it by IP rather than user-agent. Publish specific, dated, checkable detail, because domain authority and freshness drive citations. Skip llms.txt: one server-log study recorded PerplexityBot fetching robots.txt 775 times and llms.txt zero times. Blocking PerplexityBot does not remove you, since Perplexity-User ignores robots.txt.

Perplexity operates two crawlers. Most guides mention one. That gap produces a mistake running in both directions: businesses that want Perplexity traffic block the crawler that would deliver it, and businesses that want to opt out believe they have when they have not.

This article covers what Perplexity documents about itself, what independent studies measured, and where the evidence stops. Where we cannot verify something, we say so.

Perplexity built its own index

Perplexity crawls the web itself. Its engineering team reports tracking more than 200 billion unique URLs, and in April 2025 it moved retrieval in-house onto Vespa, replacing the licensed search APIs it leaned on before.

That matters for your site in one specific way. Perplexity holds its own record of your pages rather than borrowing Google’s. Ranking on Google does not carry you into Perplexity, and being invisible on Google does not lock you out. You need Perplexity’s own crawler to reach you.

Some analysts argue Perplexity still queries external search APIs alongside its own index. Perplexity’s published material emphasises the in-house index. No independent analyst has resolved how much third-party retrieval remains, so treat claims either way as unproven.

The two crawlers, and why the difference matters

Perplexity documents two user-agents.

PerplexityBot indexes your site so Perplexity can surface and link it. It respects robots.txt. Its full user-agent string reads:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible;
PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

Perplexity-User fetches a single page when a live user asks about it or pastes the URL. Perplexity’s documentation states that this one “generally ignores robots.txt rules”. Its string reads:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible;
Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

Neither crawler feeds model training, according to Perplexity’s own documentation. Any third user-agent you see quoted in an SEO blog post is not in Perplexity’s docs, so ignore it.

What this means if you want Perplexity traffic

Allow PerplexityBot. Check your robots.txt for a blanket Disallow under User-agent: *, because that catches it. Perplexity says robots.txt changes take up to 24 hours to take effect, so leave a day before you retest.

What this means if you want out

Blocking PerplexityBot stops the indexing. It does not stop Perplexity-User, which will still fetch your page when somebody asks about your business by name. A robots.txt rule gives you partial removal, not a full opt-out. Anyone who told you otherwise was working from the single-crawler version of this story.

Verify by IP, not by name

A user-agent string is text, and anybody can send it. Perplexity publishes its address ranges at perplexitybot.json and perplexity-user.json, both linked from its crawler documentation, and instructs site owners to match user-agent against IP before acting on either. Build your Cloudflare or WAF rules on that pairing.

One caution on this point. Cloudflare alleged in August 2025 that Perplexity fetched content from undeclared addresses after sites blocked it, putting the figure at three to six million daily requests. Perplexity disputed the account and attributed most of the traffic to a third-party proxy, citing a figure of 45,000 daily requests. Neither number has independent confirmation. Verify by IP because Perplexity tells you to, and keep an eye on your logs regardless.

How many sources a Perplexity answer cites

You will read a confident number for this. Two sets of figures exist and they do not agree.

Perplexity’s own publisher partnerships lead put it at four to eight sources for a typical answer. An independent analysis of 118,000 Perplexity answers by Qwairy measured an average of 21.87 citations. A separate tracker reports 5.8, ranging from 4.2 to 7.3 by query type.

Part of that gap comes from counting: a “source” and a “citation” are not the same unit, and the two camps have published no methodology that reconciles them. Until somebody does, treat any single figure you see quoted as one methodology’s answer rather than a fact about Perplexity.

The practical point survives the disagreement. Perplexity cites more sources per answer than ChatGPT, which names one or two. Your odds of appearing are better here than anywhere else, which is why the Cape Town and Johannesburg searches for Perplexity help keep arriving.

What the evidence says earns citations

The largest study we found tested 13,911 domains across 83,387 visibility records. After controlling for domain authority, the page-level tactics most agencies sell showed no significant effect on Perplexity citations. Domain authority itself dominated.

Perplexity documents that its indexing prioritises authoritative domains and undercovered topics, and that freshness feeds its ranking. It has not published the weights, so any agency quoting you a percentage for freshness is estimating.

Two things follow for a South African business.

Publishing on an undercovered topic beats competing on a saturated one. A Johannesburg engineering firm writing the only clear explanation of a local compliance requirement gives Perplexity a reason to reach for it. The tenth article about website speed gives it nothing.

Dates and specifics matter more here than on other engines, because freshness is in the documented ranking. Put publication dates on your pages, revisit them, and name figures rather than describing them.

The llms.txt question, settled

South African agencies are selling llms.txt files as Perplexity optimisation. One small study supports them and three larger ones do not.

The supporting study, by Attrifast, ran ten sites over six weeks and found a Perplexity-specific lift of 10.2 percentage points at p=0.04. Ten sites, self-run, not peer reviewed.

Against it: the 13,911-domain analysis found the effect vanished once the researchers controlled for domain authority. A 2,500-site study found sites with the file cited 1.27 times more often than their population share, which its authors described as below a meaningful statistical bar.

The finding that ends the argument came from a twelve-week server-log study across 83 sites. PerplexityBot fetched robots.txt 775 times across the panel. It fetched llms.txt zero times.

A file the crawler does not request cannot influence what the crawler does. If somebody quotes you for adding one, ask them to show you a log line where PerplexityBot fetched it.

The Publishers Program, and why you cannot join it

Perplexity launched its Publishers Program on 30 July 2024 with TIME, Der Spiegel, Fortune, Entrepreneur, The Texas Tribune and WordPress.com, then added fifteen more partners that December. Publishers earn a share of advertising revenue from queries that cite them. In August 2025 Perplexity added Comet Plus, a subscription tier that pools revenue from a reported 42.5 million dollar starting pool, keeping 20 percent and splitting 80 percent among publishers.

No self-serve route into either exists. Perplexity publishes no submission portal and no paid-inclusion mechanism. Partnership comes by negotiation, and the roster is news organisations rather than small businesses.

Any provider offering to submit your business to Perplexity or place you in a partner programme is describing something Perplexity does not sell.

What to do this week

Open your robots.txt and confirm PerplexityBot can reach the pages that describe what you sell. Wait a day.

Then ask Perplexity ten questions the way a customer would ask them. Not your company name. Questions like “who does industrial rope access in Gauteng” or “what does a five page business website cost in South Africa”. Write down which businesses it names and which sources it cites. That list is your baseline, and it tells you more than any dashboard.

Read your server logs for PerplexityBot and Perplexity-User. Match against the published IP ranges. If neither has visited, indexing is your problem and no amount of content fixes it.

Then pick one question where a competitor gets cited and you do not, and write the better answer. Include the figures, the dates and the client names you can publish.

What we cannot tell you

Perplexity publishes no South African user numbers, and neither does any neutral audience measurement body. The share figures circulating in local marketing decks come from analytics vendors selling visibility services. We have seen estimates near five percent of South African AI chatbot use, and we cannot verify them.

You may have read that a mobile operator gives South Africans free Perplexity Pro. Bharti Airtel does that in India, covering 360 million customers since July 2025. We searched for a South African or African equivalent across MTN, Vodacom, Cell C, Telkom, Rain and Airtel Africa and found none. MTN’s 2025 AI partnership is with Microsoft Copilot. If that changes, local Perplexity adoption changes with it.

Perplexity has not published how it weighs sources against each other. We test instead of guessing, and we recommend you treat any agency claiming to know the formula as guessing too.

Where this leaves you

Perplexity gives a South African business better odds than ChatGPT does, because it cites more sources per answer and rewards the undercovered topics that a specialist firm can own. It also gives you less certainty, because its ranking stays undocumented and its own citation figures conflict with independent counts.

Start with access, then baseline, then evidence. If you want us to run that baseline across Perplexity and seven other engines and hand you the raw results, our AI search visibility service sets out the scope and the price, along with a written list of what we cannot guarantee.