Myrah

    How brands measure visibility in AI search

    Myrah

    Every AI visibility number is an estimate drawn from a sample. Here is the arithmetic, and how to report it honestly.

    Key takeaways

    • Brands measure visibility in AI by running a fixed, unbranded prompt set across several AI engines on a schedule and scoring the answers.
    • Nothing logs what a model told a user, so every AI visibility metric is a sample estimate, never a count.
    • Myrah weights each engine Presence 55%, Prominence 30%, Citation 15%, averages four engines equally, and maps to 0–100.
    • Prominence uses mean reciprocal rank: positions 1, 3 and unranked give (1 + 0.33 + 0) / 3 ≈ 0.44.
    • Prompt set design is the largest source of bias, and scores are not portable between tools whose prompt sets differ.

    How brands measure visibility in AI comes down to one procedure: run a fixed, unbranded prompt set across multiple AI engines on a schedule, then score three things in the answers. How often the brand is named. How high it ranks among the brands the engine chose to name. How often its URLs are cited. Those rates get weighted per engine and averaged into one visibility score. This guide is for marketers, heads of growth and agency leads who must defend that number in a board deck or client report. Collection mechanics live in our guide to AI citation tracking; this page owns the math.

    AI visibility measurement is the practice of estimating how often and how prominently AI assistants name and cite a brand, by running a fixed prompt set across engines on a schedule and scoring the responses. It matters because no analytics log exists for AI answers, so the number is only as good as the sample behind it.

    How brands measure visibility in AI: sampling, not counting

    Web analytics counts: a session either happened or it did not, and the server saw it. Rank tracking counts too, because a results page is public and anyone can fetch it.

    An AI answer leaves no such record. If an assistant described your product to a buyer this morning, nothing you own witnessed it. No impression log, no query report, no per-answer file. Google says links in its AI features come from Google Search and need a page indexed and snippet-eligible (Google), but eligibility is not selection, and nothing there reports back to you.

    So measurement moves from counting to sampling: a fixed list of questions, asked repeatedly, on engines you choose. Every number downstream estimates a population you will never see whole. Anyone selling a precise "AI traffic" figure is overstating what is knowable.

    Four things follow: sample size matters, prompt choice is the dominant bias, run-to-run variance is real, and two vendors' scores are not comparable.

    The three metrics that matter

    Myrah scores each engine on three weighted components, published on the Myrah methodology page, which is why they can be quoted here.

    Table 1. The three AI visibility metrics, with weights, denominators, and what each cannot tell you.

    Metric Weight Unit and denominator Cannot tell you
    Presence 55% Mention rate over run responses, pulled toward a fixed per-engine baseline If the mention was favourable
    Prominence 30% MRR, 0–1, over every response in the run Anything about brands never listed
    Citation 15% % of citation-bearing responses using your URL, fading into presence as citations disappear If anyone clicked

    Presence, weighted 55%

    Presence is the mention rate: how often you were named, out of every response in the run. Named in 18 of 40 responses on one engine is 18 / 40 = 0.45, a 45% mention rate.

    Myrah does not use that raw rate directly. It pulls the rate toward a fixed per-engine baseline, so a single answer flipping in a 25-prompt set cannot swing the number the way a raw percentage would. The baseline is fixed rather than set from the current leader: a rival's good month should not drag you down while nothing about you changed. Presence carries the largest weight because it answers the first-order question — are you named at all?

    Prominence, weighted 30%

    Prominence uses mean reciprocal rank, or MRR. Reciprocal rank is 1 divided by your position among the brands the engine chose to name: first is 1/1 = 1, third is 1/3 ≈ 0.33, absent is 0. MRR is the mean of those values across every response in the run, including the ones where you never appeared.

    That denominator matters. Scoring only where you were named, or only against the competitors you happen to track, flatters the number: add a weak rival to your tracked set and your position improves without anything changing in the answers. Ranking against the brands the engine named, and taking a zero where you were absent, removes both loopholes.

    Three responses, at position 1, position 3, and unnamed:

    (1 + 0.33 + 0) / 3 = 1.33 / 3 ≈ 0.44

    MRR fits because the drop from first to second matters far more than the drop from fifth to sixth: a brand named fifth every time scores 0.20, while one named first once and missed twice scores 0.33. Being the answer occasionally beats being an afterthought consistently.

    This is also AI share of voice: your MRR beside each competitor's, same prompts, same run, read side by side in competitor benchmarking.

    Citation, weighted 15%

    Citation is how often your URL appeared as a source, counted only against responses where citations were on the table at all. The denominator is the whole point.

    Run 40 prompts. Say 22 responses displayed sources and your URL appeared in 6: the citation rate is 6 / 22 ≈ 27%, not 6 / 40 = 15%. The wrong denominator punishes you for the engine's behaviour rather than yours. Perplexity AI shows sources by default (Perplexity AI); the Gemini API returns inline citations only when google_search grounding runs (Google).

    One more property is worth knowing: as citation opportunity disappears on an engine, this component fades into presence rather than dragging you down. An engine that rarely shows sources should not penalise you for its own behaviour.

    Each component is bounded, so the per-engine raw score lands in [0,1]:

    raw = 0.55 · presence + 0.30 · prominence + 0.15 · citation

    With components of 0.45, 0.44 and 0.30, that is (0.55 × 0.45) + (0.30 × 0.44) + (0.15 × 0.30) = 0.2475 + 0.132 + 0.045 = 0.4245, which the display curve maps to roughly 42. Round it; the second decimal is noise dressed as precision.

    Why engine scores should average equally

    Myrah queries four engines, OpenAI, Google Gemini, Anthropic Claude and Perplexity AI, and averages their scores equally at 25% each.

    The alternative is traffic-weighting: give the most-used assistant the biggest share. It sounds more accurate and behaves worse. Usage shares are third-party estimates that move quarterly, so any weight buries a market-share guess inside your metric, and one vendor's new model could then swing your headline number alone. Equal weighting means a provider changing a model moves one lane of four. If a single engine matters more to your buyers, read its sub-score.

    Prompt set design is the biggest source of bias

    This is the least discussed part of AI search monitoring and the part that most determines your score. Myrah builds prompts from your niche, competitor positioning and buyer-intent phrases, in three types. Category prompts ("best X for Y") test whether you exist in the model's shortlist. Compare prompts ("X vs Y") test how you sit against named rivals. Use case prompts describe the buyer's job in their words, no vendor named.

    Branded prompts stay out. Asking an assistant "tell me about Myrah" guarantees a mention and measures nothing. A set salted with branded questions produces a flattering number that cannot fall.

    The set must also be frozen between runs. Edit three prompts and add five, and this month's score is not comparable with last month's: you measured a different thing and drew a line between unrelated points. Version it, date it, change it only at quarter boundaries.

    Size matters for a different reason than people assume. On a raw mention rate, one flipped response in a 20-prompt set moves the number five percentage points; at 40 prompts, 2.5. Myrah's presence component damps that by pulling toward a fixed baseline, but no weighting rescues a sample too small to represent the questions your buyers actually ask. Forty prompts per engine across four engines is 160 responses per run, a reasonable floor.

    Controls that make a run replicable

    Prerequisites: a versioned prompt set, a competitor list held for the quarter, and a record of the model version each engine served.

    1. Send prompts individually. Batching lets earlier answers contaminate later ones.
    2. Use no system prompt. Whatever it changes is your artifact, not the model's behaviour.
    3. Carry no conversation history. History primes the model: a brand named in turn one reappears in turn four.
    4. Append no extra instructions. "Give me five options" changes how many brands get named, and so changes prominence.
    5. Enable search tools on every engine. Retrieval produces citations, and eligibility runs through crawler access such as OAI-SearchBot for ChatGPT Search (OpenAI). A run without retrieval measures memory, not AI search visibility.
    6. Publish the model versions. Without them, nobody can separate real movement from a silent model update.
    7. Match brand names only in responses. Injecting your name into a prompt is the branded-prompt problem in disguise.

    Ask any vendor which of these they hold; the gaps in their answer are the error bars in yours. The full sequence sits on how Myrah works.

    Variance, and how to report it honestly

    Identical prompts produce different answers. Generated text is sampled rather than looked up, retrieval sets shift, and vendors update models under unchanged product names. Two runs an hour apart will disagree, which is normal.

    • Report a range or a trend line, not a single decimal. "42 to 46 across the last four runs" is honest; "44.3" is theatre.
    • Establish your noise floor. Run the frozen set twice back to back before changing anything. The gap between those runs is what your setup produces from nothing, and smaller movement is not a result. That is judgment, not a statistical test.
    • Treat small moves as noise. As a working rule, we would not build a slide around a two- or three-point move on a 40-prompt set; sustained direction across three runs is worth acting on.
    • Change one thing at a time, holding the prompt set constant.

    What good reporting looks like

    Use this as the skeleton of a monthly report.

    Table 2. Monthly AI visibility report template: each section, its source, and its common misread.

    Section What to include Where the number comes from Mistake to avoid
    Headline score One 0–100 figure and the month's range Engine scores averaged equally A decimal, or another tool's score
    Presence Mention rate per engine, with response count Named ÷ all responses A percentage without its denominator
    Prominence MRR per engine, rank against 3–5 rivals Mean of 1/position Reading low MRR as absence, not lateness
    Citation Cited rate, as % of citation-bearing responses Your URLs ÷ responses showing sources Using all responses as denominator
    Movement Trend across 3+ runs, noise floor stated Run-over-run, unchanged set Calling noise-floor movement a win
    Run metadata Prompt set version and date, model versions Prompt registry, run log Editing prompts, keeping the trend
    Actions The 3–5 ranked fixes this run produced Gaps where you were named, not cited Gaps with no owner or date

    Two rules make a report defensible: every percentage travels with its denominator, and every trend travels with its prompt set version.

    Why two tools give you different scores

    Marketers ask which platform excels in AI visibility metrics, then compare two tools on one brand and find a 20-point gap. Both can be right, because they measure different samples.

    Four things differ between AI search monitoring tools: the prompt set, the weights, the engine mix, and the display curve behind the 0–100 mapping. Change one and the score moves; change all four and the numbers have no defined relationship. A visibility score is an index, and indexes compare only against themselves.

    There is a fifth question worth asking, and almost nobody can answer it: is the method versioned, and is it ever refitted? A vendor that quietly re-tunes its weights against aggregate customer data produces a score that moves when other people's results change. Myrah's weights, baselines and curve are constants, versioned together — currently scoring v2.0.0 — changed deliberately and never refitted automatically. Ask your vendor for their version number. If there isn't one, month-over-month comparison was never safe.

    So ask which tool will tell you what it did: the prompt set and whether you can edit it, the weights, the engine list, and the model versions per run. Myrah publishes its weights, four-stage process and model versions on the methodology page, which is why those figures can appear here. That is why you use AI search monitoring tools: not for the score, but for a sample you can audit.

    Where analytics still helps, and where it stops

    Referrals from Perplexity AI, ChatGPT and other assistants arrive with identifiable referrers, so you can segment that traffic and measure what it converts at. Google's May 2026 expansion of Preferred Sources into AI Overviews and AI Mode, with link carousels and "Highly Cited" labels (Google), changes what a click looks like but keeps it countable.

    Where it stops is the answer itself. Analytics sees only the click-through minority, because most AI answers end without a visit, and no reliable public figure exists for that ratio. It cannot tell you how often you were named and not clicked, which is the gap sampling fills.

    Does AI content optimization improve search visibility?

    Sometimes, and a controlled comparison is the only way to know: freeze the prompt set, change one thing, re-run, read three runs. Google says llms.txt files, artificial chunking, AI-only rewrites and special AI schema are not requirements for its AI features, and that crawlability, technical clarity and unique, useful content are the durable work (Google), as of mid-2026. Acting on a gap is covered in our guide to improving brand visibility in AI search engines.

    Frequently asked questions

    How do brands measure visibility in AI?

    By sampling. They run a fixed set of unbranded prompts across several AI engines on a schedule, then score how often they are named, how high they rank among the brands each engine names, and how often their URLs are cited. Those rates are weighted per engine, then averaged equally.

    How many prompts do I need for a reliable AI visibility audit?

    Enough that one response cannot dominate. On a raw mention rate, one flipped response in a 20-prompt set shifts the number five percentage points. Myrah's presence component is pulled toward a fixed per-engine baseline, which damps that swing, but sample size still governs confidence. Forty per engine across four engines gives 160 responses a run.

    Why do AI visibility scores differ between tools?

    Because they measure different samples. Prompt sets, weights, engine mixes and normalization curves vary between vendors, and each moves the result. A visibility score is an index defined by its own method, so two numbers compare only if all four match.

    Can Google Analytics measure AI visibility?

    Partly. It counts referral clicks from assistants such as Perplexity and ChatGPT, with landing pages and conversions. It cannot see answers that end without a click, which is most of them. Analytics measures the traffic AI sends; sampling measures whether AI names you.

    What is share of voice in AI search?

    AI share of voice compares how often and how prominently each brand appears across the same prompt set in the same run. In practice it is your mean reciprocal rank next to each competitor's. It answers who owns the answer in your category.

    Why did my score change when I didn't change anything?

    Because AI answers are generated, not read from an index, and vendors update models behind unchanged product names. Identical prompts return different brand orders and source lists between runs. Establish your noise floor with two back-to-back runs, then ignore movement inside it.

    Conclusion

    How brands measure visibility in AI is a sampling problem wearing a counting problem's clothes. No log exists, so you build a sample, weight it, and report it with the error showing: raw = 0.55 · presence + 0.30 · prominence + 0.15 · citation, four engines averaged equally, mapped to 0–100 by a fixed curve, versioned as scoring v2.0.0. The number is only as defensible as the prompt set behind it. Write that set down with a version number this week, then run it twice back to back to find your noise floor. Read the Myrah methodology for the exact weights, or what AI visibility means.

    Sources