Method & Disclaimer

Last updated: 1 August 2026

In short

  • We measure. We do not guarantee outcomes. Nobody can promise you a place in an AI assistant's answer, including us.
  • AI answers vary and are not exactly repeatable. Ask the same question twice and you may get two different answers. That's how these systems work, not a fault in our method.
  • We report what assistants say. We do not control what they say. We have no relationship with the companies that build them.
  • Every finding is evidenced. Each report carries the exact question asked, the date and time, and the answer as it was given.

The rest of this page explains how we work and what our numbers do and do not mean. It's worth reading before you buy.

What we measure

We ask AI assistants the questions your customers ask when they're deciding what to buy, and we record what the assistants say back.

For each question we record:

We then analyse which sources are doing the work, which publications, roundups, retailers and forums the assistants keep drawing on, and what that implies about how a brand comes to be named.

That analysis is what you are paying for. The raw answers are evidence, not the product.

How the questions are chosen

Questions are written to match real buying intent in your category, not to flatter your brand. They fall into four groups:

  1. Category questions: "what's the best X for Y"
  2. Comparison questions: "X or Y, which is better for Z"
  3. Constraint questions: price, ingredient, availability or use-case limits
  4. Brand questions: direct questions about you and about named competitors

The set is proposed by us and confirmed with you before the first report. It stays fixed for at least a quarter so month-to-month figures compare like with like. Changing the questions resets the trend, so we don't change them casually.

How the work is done

Every question is written for your category by a person who has thought about how your buyers actually shop. Nothing comes off a template.

Each question is asked in its own fresh session, with memory switched off, against the assistant's standard consumer interface, the same product your customers use, not a developer API. Both halves of that matter.

Consumer interface, because API output and consumer output diverge. The consumer products layer their own retrieval, shopping data and local availability on top of the underlying model. In a Singapore reading we get product cards in SGD, local retailers, and citations to the Singapore edition of a magazine. The API returns none of that. Measuring it would tell you about a system your customers never touch. Where we later add API-based tracking for breadth, we will label it separately in your report so you always know which number came from where.

Memory off, because a signed-in account is not a neutral one. Assistants personalise answers from your past conversations, and at least one of them uses that history to shape the searches behind the answer. We have measured the same question both ways and got a different leading brand and a different set of cited sources. A reading taken inside someone's own account describes what the assistant says to them, not to a stranger deciding what to buy.

The exact conditions we measure in

A reading is only reproducible if you know how it was taken, so here is the setup, and you are welcome to copy it and check us:

The honest limitation. A clean session is not a real customer. A real shopper is signed in, has a history, and may get a different answer because of it. What we measure is what a stranger to your brand is told, which is the population you are trying to reach and the only one that can be measured the same way twice. A reading taken inside one person's account describes what the assistant says to them, and cannot be compared month to month.

Answers are then analysed rather than counted, which sources keep deciding who gets named, and what that implies for you. We use AI tooling for that analysis, in the way an accountant uses software. What it does not decide is the recommendation. Every report is read and signed off by the person you email before it reaches you, and the ranked actions are his judgment about your business rather than an output.

Sample sizes are deliberate rather than enormous. We would rather report a smaller number of observations someone has actually read than a large number nobody looked at.

Which assistants, and how often

CoreProGlobal
Buying questions75150150 per market
Assistants tracked477
Markets112, reported separately

Core covers four: ChatGPT, Google Gemini, Perplexity and Microsoft Copilot. This is not a reduced tier. Between them, those four are close to the whole consumer market for buying questions.

Pro and Global add three more: Claude, Meta AI and Grok. Meta AI carries more weight in this region than its global share suggests, because in Singapore and Malaysia it answers from inside WhatsApp and Instagram, where people already are.

We stop at seven rather than padding the number. We only track assistants your customers can actually reach from your market. One we previously considered turned out to answer only inside a retail app unavailable in Singapore, so it is not on this list. Nor are the China-market chat tools, which would raise the number without describing anyone who buys from you. An assistant your buyers cannot reach produces a figure that looks like coverage and describes nobody.

We do not track Google's AI Overviews. Google prohibits automated querying of Search and offers no API for it, so anyone claiming to track it at scale is doing it by hand or not doing it. Gemini already covers Google's ecosystem, and AI Overviews is a feature of a search results page rather than someone asking an assistant what to buy, which is the thing we measure.

Global covers two markets, and further markets are S$799 each per month. We price them separately rather than bundling because each one is a full set of questions asked again, and a plan that quietly cannot be delivered is worse than one that is priced honestly.

We name the specific assistants in your report, alongside how each one behaved for your category.

Markets available: Singapore, Malaysia and the United States. Core and Pro cover one of your choosing. Global covers two, and the third can be added.

Every question that can be run through an official API is asked three times, in three separate sessions, and reported as a frequency rather than a yes or no. "Named in 2 of 3 runs" is a more honest number than "named", and it is the only way to tell a brand that genuinely sits in an answer from one that appears occasionally. Because these systems are non-deterministic, a single observation per month would tell you very little about the underlying probability, and we would rather spend the API calls than sell you a coin flip.

The hand-run sample is asked once, because a person types it. Those questions are labelled separately in your report, so you always know which figure came from three runs and which came from one.

Reports are issued monthly. Pro and Global also receive a shorter mid-month update.

Why the numbers move, and what that means

Generative AI assistants are non-deterministic. The same question, asked twice, can produce different answers. Output varies with:

This is not a limitation we're apologising for. It's a property of the systems being measured, and any provider who tells you otherwise is misrepresenting how these products work.

What follows from it:

If you already track this somewhere else

Some of the people reading this already pay for an AI visibility tool or an agency retainer. The honest position is that those tools are not fake and their numbers are not made up. They are measuring something real. The question is whether they are measuring the thing your customers actually touch.

Three differences are worth checking against whatever you currently use. All three are testable, and you do not need us to test them.

If the choice itself is what you are weighing up, there is a longer piece on it: monitoring software, or a written report? — including the cases where the software is the better purchase.

1. Ask what interface the numbers come from

Most tools query the developer APIs, because that is the only way to run thousands of prompts cheaply. As set out above, the consumer apps are a different product built on the same models, and the gap between them is not cosmetic.

If a dashboard is showing you API output, it is describing a system your customers never open. That is not dishonest of them. It is just a different question from the one you are paying to answer.

2. Ask whether memory was switched off

A reading taken inside a signed-in account with memory on describes what the assistant says to that account, not to a stranger. We have measured it both ways; the difference is set out above.

So the question to put to your current provider is simply: was this measured in a clean session? If it was measured inside an account, the report describes what the assistant says to that account. Not to a stranger deciding what to buy.

3. Ask what you are supposed to do on Monday

A dashboard gives you a number that moved. It rarely tells you which page to write, in what words, aimed at which question. That gap is the actual work, and it is the part that does not automate: reading which sources keep deciding the answer, then deciding what to publish about it.

Every report here carries ranked actions tied to specific questions, briefs specific enough to hand to whoever writes your copy, and on Pro and Global the finished article as well. If your current provider already does that, you are well served and you should stay.

And the part that argues against us

We run fewer queries than a tool does, by a wide margin, and we always will. If what you need is breadth (thousands of prompts, tracked daily, across every market at once), a tool is the right purchase and we are not competitive on it.

We also found questions we cannot help you with. In our own measurement, the head-to-head comparison and brand reputation questions returned no citations at all: the assistant answered from what the model already held rather than anything it went and read. Publishing will not move those in the short term, and anyone who tells you otherwise is guessing. We put that in the report rather than leaving it out.

What we do not do

On Pro and Global we write content for you: one piece per month on Pro, three on Global. A piece is a single written article of roughly 800 to 1,200 words. We write it; you publish it. We never need access to your site to do this, and unused pieces do not roll over into the following month.

About competitors named in your report

Competitor brands appear in reports because AI assistants named them. We report that as a fact, attributed to the assistant and the question that produced it.

We don't evaluate competitor products, rank them on quality, or characterise them beyond what the assistant said. Where an assistant makes a claim about a competitor, we present it as the assistant's statement, not ours, and we don't adopt it as our own finding.

AI assistants sometimes produce inaccurate statements. We review outputs before including them and exclude material that appears to be a fabricated negative claim about any business.

No guarantee of outcome

Elstrand provides measurement and advice.

We do not control any AI assistant's behaviour, ranking or recommendations, and we cannot guarantee that your brand will appear in any answer, improve its position, or maintain a position it currently holds.

Recommendations in your report are advisory. Whether you act on them, how well they're executed, and whether AI systems respond to them are all outside our control.

Any statement anywhere (on this site, in an email, in a call, or in a report) should be read consistently with this page. If something we've said appears to promise an outcome, this page governs.

Change over time

AI assistants change. Models are updated, retrieval sources shift, and products are launched and withdrawn. Because of this:

Questions

If anything here is unclear, ask before you buy rather than after.

javier@elstrand.com