Last updated: August 13, 2026
Choose an AI visibility platform on how it collects its answers, then on the eight things that depend on that. Marketers hunting for the best accurate data platform for AI search optimization have usually been burned once already, by a dashboard that swung twenty points in a week with no campaign behind the move.
The cause is rarely the brand. It's the collection method and the sample size sitting underneath the number.
Quick answer: Judge an AI visibility platform on how it gets its data before you judge its features. Ask whether it queries the consumer chat interfaces, calls the model providers' APIs, or estimates from a modeled sample. Then ask how many answers it captures per prompt, per check. Those two answers decide whether everything else on the dashboard is a measurement or a guess.
But the instability isn't a suspicion any more. A variance-components study of 12,933 AI answers about brands, released in July 2026 as a single-author arXiv preprint, found that re-running the same prompt accounted for 34.8% of the variation in results, while the identity of the brand being asked about accounted for 0.7% (Category D subset, n=7,173).
Which brand you are explains almost nothing about whether you show up. How many times the tool asked explains a great deal.
AI visibility platforms collect their answers three ways, and the three don't return the same answers.
API-based tracking calls the model provider's own API and records what comes back, on a schedule the platform controls. It's repeatable and cheap to run at volume, which is why many platforms start there.
UI-scraping, sometimes called browser automation, drives a real browser session against the consumer chat interface and reads the reply a person would see. It matches the customer's experience, and it breaks whenever a vendor redesigns that interface.
Panel or sampled estimation infers visibility from a modeled subset of queries rather than capturing each answer directly. It's the cheapest to operate and the furthest from the thing you're trying to measure.
| Method | Where It Drifts | What To Ask |
|---|---|---|
| API-based tracking | The API generation can differ from the generation served in the consumer interface | "Which endpoint, and how do you check it against the live interface?" |
| UI-scraping | Interface redesigns break collection, and personalization leaks in if sessions aren't clean | "Do you query logged out, with no account, history or memory?" |
| Panel or sampled estimation | The gap between the model and the real answer set is usually undisclosed | "What share of the reported score is measured, and what share is inferred?" |
The gap between the first two has been tested in the open. An eight-week comparison of six platforms, published in March 2026 by Vismore, a vendor in the same category, found API responses and real user-facing responses diverged in roughly 23% of cases on Google AI Overviews.
On that engine, close to a quarter of the time, the API answer wasn't the answer the customer got.
Amadora.ai is our product, and the same criteria apply to it. We query the public ChatGPT, Perplexity and Gemini interfaces daily with no account, no history and no memory, because replies and cited sources diverge between the interface and the API. Claude can't be scraped, so we track it through its API as a paid add-on.
Our position is blunt: a tool measuring only through an API isn't measuring the thing your clients are asking about.
But no single method disqualifies a vendor on its own. What disqualifies one is a vendor who can't tell you which method they use, or who says "a mix" and stops there.
Profound AI, Peec AI, Otterly.AI and AthenaHQ track overlapping engine sets, and they don't all acquire their answers the same way. Ask each of them the same question and write the answers down side by side.
Nine things decide whether a platform's numbers survive contact with a client review. They're ordered by how much damage each one does when it's wrong.
Data collection method. Everything downstream is computed from whatever answers the tool managed to capture. A vendor who treats the method as an implementation detail is telling you it isn't a selling point. Ask: which engines are queried through an API, which through the interface, and which are estimated? You can read how Amadora.ai's tracking works as one worked example of a disclosed method.
Engine and platform coverage. Coverage isn't a checkbox count, because the cited-source mix differs sharply by engine. In our own tracking, Google's citations lead with YouTube, ChatGPT cites no video at all, and Perplexity surfaces Reddit (one tracked project, June 2026). A tool covering one engine hands you a content plan that's wrong for the others. Ask: which engines are in the base price, and which are add-ons?
Refresh cadence and answer volume per prompt. Cadence tells you how often the tool checks. Volume tells you how many answers each check captures, and volume is the one that decides whether the score means anything. Daily checks on a single answer per prompt produce a noisy line that reads like insight. Ask: how many answers per prompt, per check, and is that number configurable?
Citation and source-tracking granularity. You need the URL, the domain and the page type behind each mention, not a running total. Clicks are not the payoff: Pew Research Center found that just 1% of visits to a page carrying an AI summary produced a click on a source cited inside it, while ordinary result links were clicked on 8% of those visits (900 US adults, March 2025). The mention is the asset, so the tool has to show you which pages earned it. Ask: can I export every cited URL for a prompt set?
Competitor benchmarking and share of voice. A fixed list of competitor slots measures the market you already know about. Extracting every brand an answer names measures the market that exists. Across our tracked projects the count of distinct brands surfaced per project has run from 230 to 1,549 (live dashboards, May to June 2026), which is not a list anyone types in by hand. Ask: is the competitor set fixed, or extracted from the answers?
Query fan-out depth. Engines rarely answer only the prompt you tracked. They decompose it into related queries and answer those, so a platform recording your seed prompt alone is blind to most of the surface. Ask: do you capture the fan-out queries behind each answer, and can I see them?
Integrations. A visibility number nobody can join to pipeline data stays a vanity metric, because nobody funds a metric they can't tie to revenue. Ask: which joins are native, and which are a CSV somebody downloads by hand every month?
Reporting and export. Agencies live or die on this one, because a client-ready report is the billable artifact. White-labeling, scheduled sends and a raw CSV underneath are separate features, and vendors often ship some of them. Ask: can I schedule a branded report and still pull the raw rows?
Pricing transparency. A vendor who won't quote add-on prices before the call will quote them after, once you've built a reporting cycle around the tool and switching costs a quarter. Ask: which prices are published, and which ones only exist on a quote?
Two checks on any dashboard tell you whether the rest of it can be trusted: how many answers a score is computed from, and whether the prompt set behind it is padded with branded prompts. Both are usually available and rarely looked at.
Common advice is to check your visibility score weekly. While true, it's too simplistic, because a weekly check on a thin sample just resamples the noise.
A visibility score is a poll. Polls carry margins of error, and nobody reports a national result off twelve respondents, because the number would swing from one evening to the next without a single voter changing their mind.
AI visibility scores get published off samples that small.
The re-run effect is the mechanism. When one prompt asked twice returns different brand lists, a score built on a handful of answers is measuring the shuffle rather than the market.
Take one of our own tracked prompts, "AI search analytics software for marketing teams." Across a seven-day window of 23 captured answers it sits at rank 1. Across all 631 answers captured to date it sits at rank 6 (August 2026).
Same brand, same prompt, same tool. Different sample.
So what do you do with that? Ask a vendor to show you the answer count behind any score they put on a slide.
A prompt containing your brand name returns near-total visibility by construction, because the engine was handed the answer inside the question. Pad a prompt set with branded prompts and the average climbs while nothing real changes. In our experience the unbranded score is zero for roughly 90% of businesses, though that's an estimate across the projects we track rather than a census.
But strip them out before you believe anything.
Ask a vendor to filter the demo dashboard to unbranded, commercial prompts, live, on their own data. Visibility score, share of voice and average position each fail differently under a thin sample, and what each AI visibility metric measures covers them one at a time.
Check the export path before you check the logo grid, because every integration is downstream of getting the rows out.
Four joins usually matter:
But a logo grid isn't an integration list, because "works with HubSpot" can mean a native sync or a CSV you upload by hand.
Amadora.ai integrates Google Search Console, and connects through MCP to Claude Cowork, Claude Desktop, Cursor and ChatGPT Desktop. Ask each vendor which joins are native, which need a manual export, and which sit behind an API tier you'd have to upgrade into. The wiring detail lives in setting up AI visibility reporting.
Four models cover most of the category, and the prompt count is the meter in three of them.
| Model | What Meters It | What Inflates The Bill | Suits |
|---|---|---|---|
| Prompt-tier subscription | Tracked prompts per month | Adding prompts, brands or regions | Teams with a bounded, stable prompt set |
| Credit or usage-based | Actions consumed, such as audits and generated plans | Heavy analysis cycles late in the month | Teams whose workload is lumpy |
| Seat-based | Named users | Growing the team, or adding client logins | Small teams reporting internally |
| Engine add-ons | Each engine beyond the base set | Turning on every engine your buyers use | Anyone tracking more than the base three |
Most vendors combine two of these, which is where the quoted price and the real invoice part company.
Amadora.ai prices Starter at $59/month, or $49/month billed annually, for 15 prompts and one brand. Professional is $179/month, or $149/month billed annually, for 100 prompts and unlimited brands. Agency starts from $499/month, or from $399/month billed annually, at 300 prompts. ChatGPT, Perplexity and Gemini are included on every plan, while Google AI Overviews, Claude, Microsoft Copilot and Grok are add-ons from $9/month each.
But the sticker price isn't the number to model.
Model the bill at the prompt count you'll reach in month six, with every engine your buyers use switched on, and with the seats your reporting cycle needs. That figure often lands well above the plan you were quoted, and it's the only one worth comparing across vendors.
Criteria get you to a shortlist. Once you have one, see how today's platforms score on these criteria, then put the nine questions above to every vendor you book a demo with.
Report leading indicators rather than a revenue promise. Your own citation count moves with content and technical work, and your brand mention count moves with digital PR, so both respond to effort inside a single reporting cycle. Tie them to assisted pipeline where your CRM allows it, and avoid committing to a target score by a fixed date.
There isn't a portable one. Scores depend on the prompt set, the niche and how many brands the engines name, so a low score in a crowded software category can be a stronger result than a high one in a thin niche. Benchmark on your own trend line and on the competitors your tracked prompts actually surface.
Yes. One operator can run a tracked prompt set, read the citation export and brief writers from it, which is why entry plans exist at a single seat and a small prompt count. You'll want help when the work shifts from measurement to earning third-party mentions at volume.