Atlas measures whether AI platforms mention a brand, cite the brand's website, and recommend the brand. It submits a fixed set of prompts to the configured AI platforms on a schedule and scores the answers on those three outcomes. The platforms are the models behind ChatGPT, Claude, Gemini, and Perplexity, with and without web access, plus Google's AI results. The statistics on this page come from Atlas's corpus of 145,498 scored answers, current as of August 2026.
Personalization, and what Atlas measures instead
The answer a real buyer sees is personalized: platforms adjust answers using the user's location, conversation history, stored memory, and account context. No measurement tool sees another person's personalized answer, because platforms compute it per user at ask time. A tool can state this limit or ignore it. Atlas states it, and it measures the un-personalized answer: the response a platform returns with no user history attached. That baseline is the shared starting point the platforms personalize from, and it is the reading that can be compared over time and between brands, because a personalized reading would measure the test account's history along with the brand.
A prompt set can still represent two of the dimensions personalization uses: who is asking, and where they ask from. Atlas writes both into the prompts themselves, and the sections below describe how. The dimension outside a prompt set's reach, an individual user's own history and memory, stays unmeasured, and Atlas says so rather than modeling it.
Prompt sentiment influences answer sentiment
Atlas classifies each prompt's phrasing as negative ("Is X a scam?"), neutral ("Would you recommend X?"), or positive ("Why is X the best?"). In the corpus, 67.1% of answers to negative prompts included a caveat. A caveat is a warning or qualification about the brand. For neutral and positive prompts, the caveat rate was 32.0%. Average sentiment followed the same pattern: answers to negative prompts scored negative, and answers to positive prompts scored positive.
Some of that gap comes from the topics negative prompts cover rather than from the phrasing itself. For measurement, the distinction does not matter. Sentiment and caveat rates are averages over a set of answers, so a set with more negative prompts has a lower sentiment average and a higher caveat rate, whichever brand it measures.
Prompt format influences how many brands an answer mentions
Platforms mention more brands when a prompt asks for named options. "Which X should I consider?" asks for options; "How do I choose an X?" does not. In the corpus, answers to option-requesting prompts mentioned 4.55 brands on average, and answers to the other prompts mentioned 1.25. The share of answers recommending at least one brand was 22.8% for option-requesting prompts and 8.0% for the rest. Both numbers depend on the format of the prompt, so two brands' shares are comparable only when their prompt sets contain the same mix of formats.
To do this process yourself, you can replace "Atlas" with your name in the steps below. If you have questions, or if you want to talk about this methodology, please reach out to me (Tyler Einberger) on LinkedIn.
Basic prompt types & classifications
Atlas assigns each prompt a type and reports each type separately. The type states what the prompt can measure.
- Brand prompts include the brand's name: "What does X cost?", "Is X legitimate?". In the corpus, 91.6% of answers to prompts that included a brand's name mentioned that brand. When the prompt did not include the name, the mention rate was 19.6%. Averaging a 91.6% rate into results that otherwise run near 19.6% would inflate a brand's headline numbers, so Atlas reports brand prompts separately. The other 8.4% is the reason to track them at all. When a platform omits a brand from the answer to a prompt about that brand, Atlas reports the omission as an emergency finding.
- Category prompts describe a problem without including any brand's name: "best way to get X done". Atlas uses them to measure whether platforms mention the brand in answers to prompts that did not include its name.
- Comparison prompts include the brand's name and a competitor's name together: "X vs Y". Platforms tend to list the asked-about brand first in these answers because of how the answers are formatted. Atlas therefore reports positions in comparison answers separately from positions in category answers and does not average the two.
- Recommendation prompts ask for the best option for a stated need: "best X for Y". Atlas scores them on whether the brand appears among the recommended options.
Most prompts state who is asking and what they are solving
The examples above are written bare to show each type. Most tracked prompts are not bare. Buyers describe themselves and their constraints when they prompt: in one analysis of 1,827 real ChatGPT prompts, the average prompt ran 42 words. A buyer writes "I run a 20-person plumbing company and my supplier keeps missing deliveries. Who should I switch to?" and the answer to that prompt differs from the answer to "best plumbing supplier" in which brands it mentions, which pages it cites, and which caveats it raises. A prompt set written only as bare questions measures one prompting style and misses the longer, contextual style where the asker and the constraint are stated.
So most Atlas prompts state who is asking and what they are trying to solve. The who comes from the brand's recorded buyer personas. The situation comes from the same evidence the prompts are written from: customer language from calls and tickets, long Search Console queries that carry the problem in the buyer's own words, and the follow-up searches the platforms run. A therapy platform's set asks as a patient with a specific concern and as a referring clinician with a different one, rather than as one generic user.
Location is context too: for location-based businesses, Atlas scopes prompts by metro area and reports those answers by metro. And a contextualized prompt passes the same checks as a bare one. The business-fit check confirms the persona and situation fit the brand, and the framing classification scores the full phrasing, so a set with added context is still held to the 0.05 balance rule.
Where do prompts come from?
Atlas has specific algorithms that write prompts, and those algorithms work from recorded evidence about the brand being measured. Deterministic templates, dependent on the brand's business model, write the brand and comparison prompts. A generation model writes the category and recommendation prompts from the evidence supplied to it.
The following sources provide evidence and feedback on those: the brand's Search Console property supplies queries of six or more words, which preserve how people phrase the problem in full sentences. Keyword tools supply the keywords the brand's pages already rank for, with search volume. The brand's website supplies the concepts Atlas extracts from its pages. Atlas's own earlier measurements supply the follow-up searches the platforms ran while answering the brand's tracked prompts, and those searches record how each platform splits the topic into smaller searches. The brand's team supplies customer language from sales calls, support tickets, and reviews. Public forums supply the phrasings people use in communities.
Atlas links each tracked prompt to the evidence records its selection algorithm used. A prompt with no evidence can exist too; Atlas labels it ungrounded and ranks it last in the approval queue.
Implementing checks and balances in prompt selection
Atlas runs the following checks on each prompt before submitting it to any platform.
Checking prompts for business fit
The first check is business fit: can this prompt be answered sensibly about this business? Platforms will answer a templated prompt that asks a restaurant about subscription tiers and hidden fees. Those answers describe products the restaurant does not sell, and metrics computed from them describe nothing real about the business. Atlas validates each prompt against the brand's declared business model and blocks the ones that fail, recording the reason.
Checking prompts for framing balance
The second check is framing balance: whether the phrasing of the prompt set steers the answers it will be scored on. Atlas classifies each prompt's phrasing as negative, neutral, or positive with deterministic rules, so identical prompts get identical classes and anyone can audit the result. Atlas scores negative as minus one, neutral as zero, and positive as plus one, and it compares one brand to another only when each brand's prompt set averages within 0.05 of zero. The caveat rates above are the reason: averages computed from an unbalanced set differ because of the mix of prompts, not because of the brand. Atlas refuses the comparison instead of computing it.
Testing prompt sets before committing to tracking them
Atlas test-runs a prompt that passed both checks before committing to track it. The test run uses three platforms with three repeats each, and it costs about a quarter of one full tracked run across the complete platform set. It records whether the answers can be scored, which sources they cite, how many brands they mention, whether the measured brand was mentioned, and whether the repeats agree.
Each measurement supports a decision. An answer that cites only encyclopedic pages, or none, shows no sign that any brand's own content influenced it, so changes the brand publishes are unlikely to change future answers to that prompt. An answer that mentions zero brands cannot measure share. And when a prompt includes the brand's name but the answers do not mention the brand, that omission is an emergency finding on its own.
Atlas repeats each prompt because platforms return different answers to identical prompts. The corpus contains 29,344 groups where the same prompt ran on the same platform on the same day. About 1 in 20 of those groups contained both answers that mentioned the brand and answers that did not. The prompt, platform, and day were identical inside each group, so the disagreement comes from the platforms. Atlas treats a single answer as an unstable reading.
Atlas records a verdict for each tested prompt: track, rework, or reject, with reasons. A person reviews the verdict next to the test results and decides. Atlas adds prompts to the tracked set only after that approval.
Assigning prompt attributes
When a person approves a prompt, Atlas stores its prompt attributes: evidence links or its ungrounded label, its test-run results, its framing class, its prompt type, and its demand tier. Atlas computes the demand tier by breaking down entities in prompts with Named Entity Recognition, then calculates entity-based demand from the brand's own search data and ranks the prompt against the rest of that brand's set, not against absolute search volume. A niche business's prompts are therefore prioritized relative to each other rather than against consumer-scale volumes. Atlas delays the first report on a new prompt by one to two weeks so a baseline exists before anyone reads a trend. We are currently working on a way to measure entity demand alongside intents.
Built in reviews
Atlas schedules two kinds of reviews after onboarding. The first runs two to four weeks after the first tracked run. Atlas flags prompts whose early answers could not be scored, cited no sources, mentioned no brands, or disagreed across repeats, and it proposes a fix or retirement for each one. The second review repeats every quarter. Atlas drafts new candidate prompts from three kinds of change: follow-up searches the platforms now run that no tracked prompt covers, new Search Console queries, and new customer language. Atlas puts new candidates through the same two checks and the same test run as the originals.
How to select prompts like Atlas does
This is a similar process to what Atlas runs
- Collect your evidence. Export queries of six or more words from your Search Console property, list the keywords your pages already rank for, and paste in customer language from sales calls, support tickets, and reviews.
- Write down your buyer personas: who asks, and what each of them is trying to solve.
- Draft prompts in the four types above: brand, comparison, category, and recommendation. Write most of them as a person with a situation, using the wording from step 1 and the personas from step 2.
- Delete every prompt your business cannot sensibly answer. Ask of each one: could a real buyer of this business ask this?
- Label each remaining prompt negative, neutral, or positive by its phrasing. Score negative as minus one, neutral as zero, and positive as plus one, and add or remove prompts until the set averages within 0.05 of zero.
- Test-run each surviving prompt three times on three platforms. Record whether the answers can be scored, which sources they cite, how many brands they mention, whether your brand appears, and whether the repeats agree.
- Drop prompts whose test answers cite nothing or only encyclopedic pages, and prompts whose answers mention zero brands. Write down the reason each kept prompt exists.
- Run the kept prompts on a schedule, with repeats, and wait one to two weeks before reading any trend, so a baseline exists first.
- Review at two to four weeks. Fix or retire prompts whose answers could not be scored, cited no sources, mentioned no brands, or disagreed across repeats.
- Refresh quarterly. Draft new candidates from new Search Console queries and new customer language, and run them through steps 4 to 7 before tracking them.
How Atlas compares to other documented prompt selection methodologies
Before writing this section, we read the public methodology documentation of 19 AI visibility tools (August 2026). Each tool is a bit different.
- Sampling: the most commonly documented design is one run per prompt per platform per period. Atlas runs repeats on every tracked prompt and reports the rate at which the answers agree, which is how the repeat-group numbers above were measured.
- Scoring: most prompt selection methodologies don't split between prompts that include the brand's name and prompts that do not. Atlas computes it, because AI answers to prompts seeded with brand mentions expectedly returns a 91.6% mention-rate versus an average 19.6% mention-rate across category-level prompts in the Atlas corpus. A blended score is not good for what we're trying to see.
- Prompt sourcing: several tools ground their measured prompts in large keyword or prompt corpora from clickstream providers, and others generate prompts per brand at onboarding. Atlas grounds each brand's prompts in that brand's own recorded evidence (its Search Console queries, its customers' language, and the fan-out searches captured from its own tracked answers) and validates each prompt against the brand's declared business model, because a 2026 incident in our own history showed what an unvalidated template produces: it asked restaurants about subscription tiers, and the answers scored products the businesses do not sell.
- Prompt maintenance: a handful of methodologies document a review cadence for the prompt set. Atlas reviews the prompt sets at two weeks, then refreshes it quarterly from new search data and new customer language, and keeps records of why each prompt exists.
- Prompt style: we are not aware of a methodology that features a first-person, persona-framed example prompt. Some methodologies publish bare keyword-style prompt examples, and one specifically prescribes two-to-five-word prompts deliberately, as a tracking-stability choice. Bare prompts also increase the measured number: a 37,804-response study found keyword-style prompts produce up to 25% more brand mentions than persona-framed prompts. Atlas writes most prompts as a person with a situation anyway, because that is closer to what buyers type, and we enjoy a less inflated analysis that it seems to produce.
The most important part about methodology (in my opinion) is knowing metrics are only comparable when the prompts and collection behind them are controlled.
