Measure AI recommendations with a fixed prompt set and consistent test conditions. Record mentions, citations, accuracy, sentiment, competitor inclusion and source URLs for each engine. Then connect those observations to referral sessions, engaged visits, enquiries and assisted conversions. Keep manual prompt results separate from analytics, because no tool can provide a universal AI ranking.
A favourable screenshot is not a measurement system.
ChatGPT can name a business in one answer and omit it after a small prompt change. Gemini can use different sources under a different account or location. Perplexity can cite the website without recommending the company.
Measure the pattern, the accuracy and the business result.
Separate three types of evidence
AI visibility reports often combine unrelated numbers into one score. Keep three evidence layers separate.
1. Manual answer observations
These show what an engine returned under recorded conditions.
Examples:
- The business was named.
- The website was cited.
- A competitor appeared first.
- The answer stated the wrong city.
2. Platform and website data
These show discoverability and visits.
Examples:
- Google generative AI impressions.
- ChatGPT referral sessions.
- Perplexity referral sessions.
- Crawler requests in server logs.
- Landing-page engagement.
3. Business outcomes
These show whether visibility helped the company.
Examples:
- Qualified enquiries.
- Calls.
- Proposal requests.
- Assisted conversions.
- Revenue linked to a known lead source.
A crawler visit is not a recommendation. A recommendation is not a lead. A lead is not revenue.
Build a fixed prompt set
Use twenty prompts for a stable small-business baseline.
Divide them into:
- Four category prompts.
- Four problem prompts.
- Four comparison prompts.
- Four evidence prompts.
- Four branded accuracy prompts.
Do not place the company name inside discovery prompts. That tests whether the engine selects the business without being told where to look.
Write prompts with enough context to produce a useful answer.
Weak:
Best company near me?
Stronger:
Which Pretoria web design studios publish pricing and let a small-business client own the domain and website files?
The second prompt states the location, category and decision criteria.
Fix the test conditions
Record the conditions with every answer.
| Field | Example |
|---|---|
| Test ID | P-07 |
| Date and time | 2026-08-06 10:00 SAST |
| Engine | ChatGPT |
| Product or mode | Search |
| Account state | Logged out |
| Device | Desktop browser |
| Location context | Pretoria, South Africa stated in prompt |
| Exact prompt | Saved without edits |
| Web retrieval | Active, inactive or unclear |
| Clean conversation | Yes or no |
| Answer capture | Screenshot or export path |
| Citation capture | Full URLs |
Use a new conversation for each prompt. Prior messages can change the answer.
Repeat the same setup in each cycle where possible. Perfect control is impossible, but sloppy testing creates avoidable noise.
The ten metrics that matter
1. Prompt coverage
Prompt coverage shows how much of the planned test ran successfully.
completed prompts / planned prompts
A report based on 12 of 20 prompts should say 60% coverage. Do not present it as a complete baseline.
2. Mention share
Mention share measures how often the business appears.
prompts that named the business / completed discovery prompts
Keep branded prompts out of the discovery rate.
3. Citation share
Citation share measures how often the company’s own website supports the answer.
prompts that cited the website / completed prompts
Track third-party citations separately. A directory can drive a mention while the website remains absent.
4. Accuracy
Score the important business facts:
- Service.
- Location.
- Audience.
- Price or pricing model.
- Founder or team.
- Contact route.
Use:
- 2: correct.
- 1: incomplete.
- 0: wrong or unsupported.
A mention with a wrong service can create low-quality enquiries.
5. Sentiment and framing
Record whether the answer presents the business positively, neutrally or negatively.
Do not treat every positive adjective as a win. Note the reason the engine gave and whether a citation supports it.
6. Competitor inclusion
Record every named competitor and the criteria used to include them.
This shows which businesses repeatedly own a topic, city or proof type.
Do not turn the report into an attack list. Use it to identify missing evidence.
7. Referred sessions
Track identifiable visits from AI products.
OpenAI says ChatGPT referral URLs include utm_source=chatgpt.com, which makes those visits easier to filter in analytics.
Other products can send a standard referrer, strip it or open the site in a way that appears as direct traffic.
8. Engaged sessions
A cited page can attract curiosity without commercial fit.
Review:
- Engagement time.
- Pages viewed.
- Scroll depth where available.
- Service-page visits.
- Contact-page visits.
9. Enquiries and assisted conversions
Record form submissions, calls, WhatsApp actions and proposal requests from known or self-reported AI discovery.
Ask:
Where did you first hear about us?
Allow more than one source when the journey involved AI search, Google and a referral.
10. Test volatility
Volatility shows how often the result changes across repeated runs.
prompts with a changed mention or citation result / repeated prompts
A business that appears in one of three runs has a different position from a business that appears in all three.
The scorecard
Use one row per prompt and engine.
| Prompt ID | Engine | Mention | Website citation | Other citation | Accuracy /6 | Sentiment | Competitors | Referral visits | Enquiries | Notes |
|---|---|---|---|---|---|---|---|---|---|---|
| P-01 | ChatGPT | 1 | 0 | Directory | 5 | Neutral | A, B | 3 | 0 | Correct service, old phone source |
| P-01 | Gemini | 0 | 0 | Competitor site | 0 | N/A | B, C | 0 | 0 | Business absent |
| P-01 | Perplexity | 1 | 1 | Review site | 6 | Positive | A | 7 | 1 | Website cited for ownership policy |
The table above shows the structure, not IDJOY performance data.
Do not publish example rows as measured results.
Measure each engine according to its available data
ChatGPT
Use manual prompt tests for answer behaviour.
Filter analytics for utm_source=chatgpt.com and known ChatGPT referrals. Keep visits separate from the manual mention rate.
Check that OAI-SearchBot can access pages intended for search discovery.
Google AI features and Gemini
Google says AI Overviews and AI Mode use its Search systems.
On 3 June 2026, Google announced dedicated Generative AI performance reports in Search Console for a subset of websites. The report can show impressions, pages, countries, devices and dates.
Use it when available. State when the property does not yet have access.
Gemini product answers can differ from Google Search AI features. Test the named product instead of treating all Google AI output as one engine.
Perplexity
Use manual prompt tests and citation capture.
Check PerplexityBot access and server logs where available. Crawler requests show access, not answer inclusion.
Track Perplexity referral domains in analytics, then connect the landing page to actions.
Diagnose before changing the website
A weak result can have several causes.
| Observation | Likely gap | First check |
|---|---|---|
| No mention and no citation | Access or topic coverage | robots.txt, indexing and source page |
| Competitor cited for your service | Evidence gap | competitor’s cited page and criteria |
| Business named with wrong facts | Entity inconsistency | cited source and public profiles |
| Website cited but business not selected | Commercial clarity | service scope, fit and proof |
| Visits without enquiries | Conversion gap | landing page, CTA and offer |
| Strong branded results only | Discovery gap | non-branded content and corroboration |
| Results change every run | Volatility | test conditions and repeated samples |
Do not prescribe schema for every problem. Fix the observed gap.
Build a monthly report
A useful report answers six questions.
- Which prompts ran?
- Where did the business appear?
- Which sources influenced the answers?
- Which facts were wrong?
- Did anyone visit and enquire?
- What will change next?
Include:
- Test coverage.
- Mention and citation share by engine.
- Accuracy issues.
- Competitor patterns.
- New or lost citations.
- Google generative AI visibility where available.
- Referral sessions and engagement.
- Enquiries and lead quality.
- Changes made since the previous cycle.
- Limitations.
Keep the raw answer captures so another person can inspect the claims.
Avoid a single invented AI score
A single number can hide a weak method.
A provider may combine branded mentions, crawler requests, directory citations and traffic into a 78% “AI visibility score.” The number looks precise while the buyer cannot reproduce it.
Composite scores can help internal reporting when the formula remains visible. Publish the components and weights.
For example:
30% discovery mention share
20% website citation share
20% factual accuracy
15% qualified referral engagement
15% enquiry contribution
The business should still see the raw rows.
How often should you test?
Monthly testing suits most South African service businesses.
Test more often when:
- The business launches a major service.
- A rebrand changes the entity.
- The website moves domains.
- A public error spreads across sources.
- An engine changes a product or reporting method.
Avoid daily checking. Frequent manual searches encourage teams to react to noise.
What a credible agency should disclose
Ask the provider for:
- The exact prompts.
- The engines and modes.
- The account and location conditions.
- Raw answer captures.
- Citation URLs.
- The scoring formula.
- The difference between observation and analytics.
- Missing data.
- Failed prompts.
- The result the provider cannot guarantee.
A report that shows only favourable screenshots is marketing material.
IDJOY’s measurement position
IDJOY uses fixed prompts, raw answer evidence, crawler checks, analytics and enquiry data where the client can provide access.
The report separates:
- Being crawlable.
- Being cited.
- Being named.
- Being described accurately.
- Receiving a visit.
- Receiving a qualified enquiry.
IDJOY does not sell a universal ranking or guaranteed recommendation.
The useful outcome is a record the business can inspect and a clear next action tied to the evidence.
Read how to improve AI search visibility in South Africa before choosing the work, then review the AI search visibility service for the implementation scope.