Firemark reports the answers it collected, counted as measured. Every answer is scored, uncertain cases are checked again, and fixed rules decide what we are willing to show. Every run is a sample, and this page says how large that sample is and where the method stops.
A run asks every prompt in your basket on every engine in your plan, scores each answer and publishes the results. Answers per run are prompts x engines, so 100 prompts on 4 engines is up to 400 answers.
Your basket of prompts, each with a funnel stage, is the fixed set of questions.
Before the run, we check that each engine is answering. An engine that is not is shown as short.
Each prompt is asked on each engine from a US location, in English.
Every answer is scored for who it names, who it links and where.
The run is published, and counts, ranks and claims are worked out from it.
The engines we ask
ChatGPT (web). Asked from a US location, in English.
Perplexity. Asked from a US location, in English.
Google AI Overviews. Asked from a US location, in English.
Google AI Mode. Asked from a US location, in English.
Gemini. Asked from a US location, in English.
The schedule
Your first run starts as soon as you confirm your basket, and most brands see first results within about an hour. After that you choose daily, weekly or monthly, up to your plan's fastest cadence, and scheduled runs start overnight, US Eastern time.
If an engine does not answer enough of your prompts, we retry before the run is published as short. A failed answer is never counted as your brand being absent.
What do we record for each answer?
For every answer we record whether your brand is named, whether your site is linked, which funnel stage the prompt belongs to and which sites the answer cites. Named and linked are counted separately.
Named
Your brand's name or one of its aliases appears in the text of the answer. A name found only inside a list of sources does not count.
Linked
The answer's sources include your own site. A link alone never counts as named. The two give four verdicts: named and linked, named, linked but not named, and absent.
Funnel stage
Every prompt is tagged top, middle or bottom of the funnel, so you can see where in the buying journey an engine names you.
Sources
The sites an answer cites. We count the answers that cite a site at least once, split by answers where you are named and answers where a competitor is named and you are not.
Position and sentiment
Which third of the answer holds your first mention, and whether the mention reads positive, neutral or negative. Sentiment is recorded only when you are named.
Snippet
The sentence that holds the mention, up to 500 characters. Full answers can be read in the app for 13 months.
How are answers scored?
A scoring model decides whether each brand name in an answer is used as the name of that company or product. Checks run in layers: simple checks first, a model for the rest, and a second review for names it cannot settle. A name still undecided counts as not named.
Whether an answer links you, where your first mention falls and the snippet are worked out by code from the answer and its sources. The scorer is a model, so it can be wrong on an edge case, which is why uncertain names get a second look.
How is an AI visibility score calculated?
Firemark reports counts instead of one blended score: how many answers name you, how many link your site, and where you rank among the competitors you track.
Each answer gets a verdict. Using the scoring rules above, your brand is named or not named, and linked or not linked. The two are counted separately.
Your share is a count over a base. It is the answers that name you divided by the answers collected, for one engine or for all engines in your plan. A share is shown only when its base is large enough to mean something.
Movement needs repeat runs. Up or down is reported only after several comparable runs and a change large enough to be more than noise.
There is no universal good score. The figure depends on your prompts, your market and your competitors, so read your share against theirs and across several runs.
What is AI share of voice?
Share of voice is the same count for every brand you track: the answers that name each brand, out of the answers collected. Your rank places you among those brands. A competitor named in very few answers of a run is shown without a rank for that run.
How big is a sample, and why do results vary?
Each engine contributes one answer per prompt per run, so for one engine your sample is your prompt count. Across engines it is prompts x engines: up to 75 answers on Lite, 400 on Starter, 1,000 on Growth and 2,500 on Pro.
AI engines do not give the same answer to the same question every time, and a small basket moves more than a large one. The Engines screen draws a 95% range on each engine's own base, so you can see how far a share could move from the sample alone. We do not run a significance test, and we do not put a range around a base that is too small to support one.
What a 40% share looks like at each basket sizeArithmetic only (a 95% range on one engine's base), shown to illustrate sample size. It is not data from a real brand.
Plan's prompts
Named in
95% range
25 prompts (Lite)
10 of 25 answers
23% to 59%
100 prompts (Starter)
40 of 100 answers
31% to 50%
200 prompts (Growth)
80 of 200 answers
33% to 47%
500 prompts (Pro)
200 of 500 answers
36% to 44%
A range like this covers sampling alone. It does not cover an engine changing how it answers, which is why a change of collection method restarts the comparison. See plans and answers per run.
What are the limits of this method?
Where and how we ask
United States and English only. Every engine is asked from a US location in English.
Answers are collected from a logged-out session. A signed-in person with history can get a different answer.
Claude is not tracked yet. It is added to Pro when it ships.
What you get back
Results are samples of what engines say when asked your prompts. Another phrasing can get another answer.
MCP and CSV carry snippets of up to 500 characters and no full answers. Full answers can be read in the app for 13 months.
Counts, verdicts and snippets are kept for the life of your account.
What do we not track?
Countries and languages other than the United States and English.
Personalised or logged-in answers.
Claude, until it ships.
Clicks, visits or sales that follow an AI answer. We record what answers say, and what people do next is outside the data.
Whether what an answer says about you is accurate. We record named, linked, position and sentiment.
We review this page against how the product works whenever the method changes.
AI visibility accuracy FAQ
How accurate is AI visibility tracking?
Each answer is scored by a model, and uncertain answers are checked a second time. Firemark reports counts as measured, shows a percentage only when the sample is large enough, and reports a movement only after repeated runs. See the limits of the method on this page.
Which countries and languages does it cover?
The United States and English only. Every engine is asked from a US location in English.
What counts as named, and what counts as linked?
Named means the brand's name appears in the text of the answer. Linked means the answer's sources include your own site. They are counted separately, so an answer can be named and linked, named only, linked only, or absent. A name that appears only inside a list of sources does not count as named.
Why do results change from one run to the next?
AI engines do not give the same answer to the same question every time, so a run is a sample. We report counts as measured, draw a 95% range per engine on the Engines screen, and call a movement only after three comparable runs and a change of 5 points or more.
See the numbers for your brand
Start tracking, or open the demo to see a full report first.