Citation measurement protocol
How we measure whether an answer engine cites a source
Agents·Index runs a standing measurement of which sources four answer engines cite when asked about the agents in this index. This page is the method, published before any result, so that a stranger can reproduce a run from this page alone and check our arithmetic rather than take our word for it.
On this page
What counts as a citation
One of our URLs appears in the answer's source annotations. That is the whole definition, and it is deliberately stricter than the one most reports use.
A brand name appearing in the prose does not count. A mention a reader cannot click sends no traffic and earns no link, and counting it would let us report a rising number while nothing about our position had changed. We would rather publish a smaller number we can defend.
A redirector is not a source. Some engines return annotation URLs that point at their own redirect infrastructure rather than at the publisher. We resolve those to the real host where the response carries it, and discard the annotation where it does not. Left uncorrected, this single detail would make the redirecting engine's own parent company the most-cited domain in the dataset.
The engines, and the exact models
Four engines, queried through a metered provider API rather than a browser session. The model is part of the protocol, not an implementation detail. A different tier of the same engine can cite differently, so a reproduction that does not match these is measuring something else.
- ChatGPT:
gpt-4o-mini - Claude:
claude-haiku-4-5 - Perplexity:
sonar - Gemini:
gemini-2.5-flash
These are the cheap current-generation tiers of each engine, chosen on purpose: this measures whether a source gets cited, not how good the answer is. Cost is the reason the cadence differs per engine below, and cost is a methodological fact rather than an excuse: a protocol nobody can afford to run weekly is a protocol that runs once.
The question set is derived, not written
Nobody hand-picks the questions. They are regenerated before every pass from three sources: the published agents, the live sitemaps, and external search-demand data. So publishing a page adds its question automatically, no question can quietly outlive the page it was asked about, and the part of the set that measures the wider category does not depend on what we happen to have published.
The set has five tiers, because the shapes differ in cost and in how fast they move:
- Vendor. One question per published agent. The widest tier and the one that matches real demand.
- Vendor (hot). The same questions, restricted to the agents that engines actually fetched from us recently. Membership is regenerated from the retrieval log, so this tier follows demand instead of following a list somebody typed.
- Filtered. One question per collection page, read off its own slug: the attribute-intersection shape.
- Head. A small set of category head terms. We do not expect to win these and keep them anyway: the source list behind a lost answer names who holds the slot, which makes a zero here competitive intelligence rather than a failing lane.
- Category. The one tier that is not derived from our own corpus. Its questions are the highest-demand queries in the category, taken from external search-volume data rather than from our slugs, so the set does not move when our index does and a question survives whether or not we have a page for it. We expect to be absent from most of these answers. That is the point: the value is in the full source list, which names who was there instead.
How the category is bounded, since the noun is ambiguous. Demand data for a term like “agents” is full of other industries: estate agents, voice-over talent agents, marketing and travel agencies. We filter with a published pair of rules: a keyword must match an in-category vocabulary and must not match a known other-industry term. Both halves are needed. A blocklist alone keeps a term like “digital marketing agents”, which names no excluded word and is nonetheless about human agencies; an allowlist alone drifts as vocabulary grows. The pair also keeps a distinction either one alone destroys: “AI voice agents” is in scope and “voice-over agents” is not. Terms we cannot classify are dropped rather than guessed, and near-duplicates that differ only by word order or plural are collapsed to one question, so no answer is bought, or counted, twice.
Repetition, cadence and dates
One call per question per engine per pass. Not three, not ten. Answer engines are non-deterministic, so a single pass is a sample of one and must not be read as a reading — only the trend across passes is interpretable. We state this plainly because a protocol that quietly implies repetition it does not perform is worse than one that admits a single shot.
Cadence is weekly or monthly depending on the tier, staggered by engine so the metered spend stays flat: the cheapest engine runs weekly across the wide inventory and the category set, and the expensive engines run monthly against the narrower tiers — the hotagents, the head terms, and the category set on the other three engines. That last pass is the one that lets a single question be compared across four engines on the same day. Every row is stamped with the date it was collected, and a row is never back-dated or re-stamped when it is re-read.
Nulls are kept. Partial passes are refused.
Every question asked records a row, including the ones where nobody was cited.A dataset that only retains hits cannot produce a rate, and a rate is the only thing worth publishing. The zeros are the denominator.
A pass that cannot complete is discarded rather than recorded. Every run has a hard ceiling on metered calls; when a run would exceed it, the whole pass aborts and fails loudly instead of writing the questions it managed to reach. This matters more than it sounds: a truncated pass looks exactly like a completed one in the data, and it would silently turn a measurement into a sample nobody declared. We would rather have a missing week than a week that lies.
Only live properties are measured. A retired property is never asked about, so a discontinued page cannot inflate a coverage figure by remaining in the question set.
How this can be wrong
A method that lists no way to be wrong is marketing. Every one of these is real, current, and would change how a reading should be interpreted:
- Single-shot sampling. One call per question per pass against a stochastic system. A change between two passes is not evidence of anything on its own. Read trends over many passes, never a delta between two.
- Two of the tiers measure us, one measures the category. The vendor, hot and filtered tiers are derived from our own corpus, so a reading from them describes visibility on the topics we publish about and nothing wider. Only the category tier, whose questions come from external demand data, supports a category-wide statement. Any figure we publish says which tier it came from; a number without a tier is not a number we produced.
- The category boundary is a judgement, and it is ours. The in-category and excluded-industry vocabularies above are written by us and applied automatically. They are deliberately conservative — an ambiguous term is dropped rather than guessed — which means the category set under-covers rather than over-covers, and a vendor whose buyers use vocabulary we did not anticipate will be under-represented. The rules are published so the bound can be argued with, which is the only honest way to hold a boundary somebody has to draw.
- A cheap model tier. Results describe the model named above. A frontier tier of the same engine may retrieve and cite differently.
- An API, not a person. We measure what the provider API returns. A logged-in human, in a different country, with personalisation and memory on, may see different sources for the same question.
- The engines change silently. Retrieval behaviour shifts without announcement, and published studies have watched a single engine's citation rate for a major source move several-fold inside a few weeks. A fall in our numbers may be the engine and not the source; so may a rise.
- A strict definition cuts both ways. Counting only clickable source annotations means we under-report influence: being named in the prose without a link is worth something we do not capture.
What would make us retract a reading
Stated in advance, so the test cannot be chosen after the result: we withdraw or re-run a published reading if the pass behind it turns out to have been truncated, if an engine's endpoint or model changed inside the collection window, if the annotation parsing counted a redirector as a publisher, or if the question set was regenerated mid-pass so that different engines were asked different things. Any correction is published on the edition it affects, with the original left visible and marked, never silently edited.
The data
No edition is published yet, and this page shipping first is the point. The method is committed to before the results exist, so the results cannot shape it. When the first edition is published it will carry the dataset, the collection dates and a link back to this page — and this page will link forward to it.
A time series only gets stronger by existing longer, and it cannot be back-filled: a measurement started later can never recover the weeks it did not take. That is the reason to run this before it is interesting.
Independence
Agents·Index sells placement and never a verdict, and this measurement is not for sale in either direction: no listed party can pay to be measured, to be excluded, or to have a reading changed. Where a agent we measure is also a commercial relationship, that is disclosed on the page where it appears.
This is a different document from how we rank, which covers how a agent earns a page here. This one covers only how the citation measurement is taken. Found an error in the method? Tell us — a protocol that cannot be corrected in public is not a protocol.