WhatsApp us
GEO & AI Search

How to Measure Your Brand’s Visibility in ChatGPT and AI Search

Build a repeatable AI visibility report with clear prompt samples, recommendation and citation rates, failed-run handling and current Google and Bing data.

A modern workspace showing hands analyzing marketing data on a laptop and paper documents.
Photo: Kindel Media / Pexels
Share

A quicker read

Summarize with AI

Five key points. Three practical next steps. Choose your assistant, then paste the copied article into your new chat.

Preview the summary request

You choose what to send. AI summaries can miss details; use the full article for context.

Measure AI visibility by collecting answers to a defined set of customer questions, recording mentions, recommendations and citations separately, and repeating the same method over time. Report which platforms, languages and questions you tested, how many answers you obtained and how you handled failures. Use native platform reports alongside this sample, keeping their metrics separate.

A percentage without that context is difficult to interpret. “Visible in 40% of answers” might describe a useful set of unbranded buying questions, or a flattering set that repeatedly names your company. Neither is automatically your share of Dubai's AI-search audience.

Define what you want to observe

Use a small set of clearly named measures before adding a composite score:

Brand mention: the answer names your business, including an agreed spelling or alias. Count it once per answer, even if the name appears repeatedly.

Recommendation: the answer presents your business as an option for the customer's stated need. A passing reference, warning or explicit rejection should not count as a positive recommendation.

Own-domain citation: the answer attaches a supporting citation to a URL on your agreed website domain. A directory mentioning your business is an external-source citation, not a citation to your website.

Factual accuracy: whether the answer represents important business facts correctly. Record inaccurate services, locations, prices or contact details separately from presence. An incorrect recommendation can still be visible, but it needs attention rather than a success label.

Write these definitions down and use them consistently. Decide how to handle trading names, branches, parent companies and acquired brands before comparing results. Keep unclear cases available for manual review.

Choose questions from real buying decisions

Build the starting list from enquiries, sales conversations, service choices and comparisons your customers genuinely make. Group it by purpose, such as finding a provider, comparing alternatives or checking suitability for a specific need.

For a Dubai training provider, relevant questions could distinguish corporate workshops, individual evening classes and courses delivered at a client's premises. Those purchases need different evidence. Testing only “best training company in Dubai” would miss much of that distinction.

Maintain two separate sets:

  • Discovery questions do not name your brand. They test whether it appears as an option without being suggested.
  • Branded accuracy questions explicitly ask about your company. They test recognition and representation, not unsolicited discovery.

Keep a stable core for comparisons. Put new experimental questions in a separate, dated group. If you change the core, preserve the previous version and explain the change instead of comparing the new score as though nothing happened.

Use Arabic and English sets where both matter commercially. Have someone competent in each language review whether the questions express equivalent buyer needs. Separate the results so a change in language mix does not masquerade as improved visibility.

Record the conditions, not just the answer

For each observation, save:

  • Prompt ID and exact wording, including stated service area or location.
  • Platform, consumer app or API, and visible model or mode where available.
  • Date, time and language.
  • Known account, conversation and location conditions that could affect the result.
  • Whether the intended search behaviour occurred, where observable.
  • The complete answer, supporting URLs and your coding of each outcome.
  • Any collection error, missing result or retry.

A new conversation helps avoid carrying over earlier questions, but it does not by itself prove that account memory or location context has disappeared. OpenAI documents that memory and location can influence ChatGPT Search. ChatGPT Search context.

Do not label API responses as consumer-app observations. A monitoring tool may use an API, browser collection or another method; record which one it actually uses. Comparability depends on the collection method, not simply the model name printed on a dashboard.

Store only what is needed for the study. Remove unrelated personal information from answer captures before circulating reports.

Use denominators that readers can inspect

Here is a hypothetical example, not a Lunasol result or a recommended minimum sample size. One platform and mode are tested with 20 unique unbranded questions, twice each on a preplanned schedule.

That creates 40 planned observations. Four technical collection failures leave 36 usable answers. Twelve mention the brand, nine recommend it and six cite its website.

Measure Calculation in this example Result
Collection completion 36 usable answers ÷ 40 planned observations 90%
Brand mention rate 12 answers mentioning the brand ÷ 36 usable answers 33.3%
Recommendation rate 9 answers recommending the brand ÷ 36 usable answers 25%
Own-domain citation rate 6 answers citing the website ÷ 36 usable answers 16.7%

These are proposed sample metrics, not official cross-platform definitions. The three presence outcomes can overlap; do not add their percentages. There are 20 unique questions, not 40 unique questions or 36 different customers. Repeated answers are also not necessarily statistically independent.

The recommendation rate describes the 36 usable answers. The 90% completion rate exposes the missing observations, whose outcomes are unknown. Neither number establishes how often real customers would see the brand.

Decide failure handling before collecting results

A valid answer that does not recommend your business stays in the denominator. A response saying that it cannot identify a suitable provider can also be a valid negative observation. A tool timeout or missing capture is a collection failure, with its reason recorded separately.

Set a consistent retry rule in advance. Do not keep rerunning unfavourable answers until the company appears. If the protocol specifically requires web search and an answer does not use it, preserve that observation in a separate category rather than quietly discarding it.

For Google AI Overviews, a successfully observed search with no Overview is a non-trigger, not a broken collection. If you calculate inclusion only among searches that triggered an Overview, disclose that denominator and the number of non-triggers as well.

Compare like with like across time and competitors

Compare competitors on the same valid answers and with the same coding rules. Count each brand at most once per answer for these presence rates. Include every named competitor in the agreed comparison, rather than replacing strong performers with easier ones between reports.

Keep results separated by platform, language and buyer-intent group. A blended average can rise simply because more easy branded questions or more favourable platforms were tested.

Where completion differs between periods, inspect the matched questions and runs as well as the headline rate. Losing difficult prompts from the denominator can make performance look better without any change in the answers you actually collected.

Choose a collection schedule your team can sustain. Repeat observations help reveal variability, but there is no universal prompt count or frequency that makes a handpicked set representative of every customer. Preserve a change log for published content, tracking methods, prompt versions and visible platform changes.

Add first-party platform evidence without mixing units

Google's Generative AI performance report for Search reports link impressions in AI Overviews and AI Mode, with page, country, device and date views. It does not measure every unlinked brand mention or every recommendation. Property totals and page totals can differ because of aggregation; do not assume summing page rows reproduces the property total. Google's report documentation.

These are platform-observed impressions. They should sit beside your selected-prompt results, not be added to the number of test answers or averaged with a sample recommendation percentage.

Microsoft's AI Performance reporting covers citations across supported Copilot, Bing and partner experiences, including cited pages and sampled grounding-query information. Its scope is not a census of all AI assistants. Bing's AI Performance overview.

Microsoft also offers a preview Citation Share measure for a grounding query: the site's citations as a proportion of all sites' citations for that query. Microsoft explicitly distinguishes it from traffic share, quality scores and a competitor-domain list. It is a different denominator from the answer-presence rates above. Bing's expanded AI visibility reporting.

Make the report useful for the next decision

A practical report should let the reader see the sample, completion, separate outcomes and changes for comparable groups. Include examples of newly observed recommendations, lost citations and material factual errors, with the underlying answers available for checking.

Then identify what needs investigation. More inaccurate mentions may justify correcting business information. Citation growth on an unrelated topic may have little commercial value. A movement after a content update is worth examining, but timing alone does not prove the update caused it.

If commissioning Lunasol's GEO service in Dubai, agree the priority questions and reporting definitions with the team before the first baseline. The useful connection is between observed gaps and the content or business-information work commissioned to address them.

Keep visits, qualified enquiries and sales in a separate acquisition report. Visibility measurement tells you how your business appeared in the evidence collected; it does not substitute for measuring what customers did next.

WhatsApp us