Reference
Methodology
How the four metrics in this product are computed, at the definition versions in force. This methodology is generated from the metric definitions in code, and any disagreement between the two is a defect.
- Metric definitions
- mention_rate.v1,qualified_recommendation_rate.v1,citation_rate.v1,share_of_recommendation.v1
- Share of recommendation points
- sor-points.v1
- Revised
- 29 September 2026
What the figures are
These are four separate measurements. They are not combined into a single score.
Each figure is computed for one tracked entity, such as the brand or a competitor, over one prompt cohort and one period. It is stored with its numerator, its denominator, its universe, its sample size, its interval with the method that produced it, and the version of the definition it was computed under.
A stored figure is never edited: computing again writes a new snapshot beside the old one. Each definition carries its own version, and a change to any word, rule or input of a definition is a new version of it.
Which answers count
An eligible answer is an answer collected through the API series for the project and prompt cohort a snapshot names, observed inside the snapshot's period. The period is half-open: it includes its first moment and excludes its last, so two adjacent periods never count the same answer twice. The number of eligible answers is the universe every figure reports.
An eligible answer is covered when its current extraction, the highest version, has the status classified or model skipped and was produced by the snapshot's extraction version set: the extractor, citation rules and mention rules versions, and for a classified answer the extraction prompt and rubric versions as well. An answer not yet extracted, extracted by an earlier version, or whose extraction did not succeed is not covered, and the snapshot records its coverage as partial. One snapshot therefore never mixes the output of two extractors.
Manual observations, recorded by a person who read an answer, are a separate series. They are never counted in these figures.
Analyst corrections are not applied to the figures. Each figure reads the extraction as the extractor stored it, and each snapshot counts the covered answers that carry a correction.
Where the set a figure is taken over is empty, the figure is not measured. Where an input the figure needs is not recorded, the figure is not measurable. Neither is ever stored as a zero, and each names the input that was missing.
The four metrics
Each metric below is read directly from its definition in code, in the order every surface lists them.
Mention rate
- Definition version
- mention_rate.v1
- Unit
- A proportion: a count of answers over a count of answers, from zero to one.
- Estimand
- The proportion of the cohort's API answers in the period that name the entity.
- Numerator
- The covered answers that name the entity at least once at a stored mention span.
- Denominator
- The number of the covered answers that track the entity (an answer is covered when its current extraction was produced by the snapshot's extraction version set, and it tracks the entity when that extraction holds a role for it).
- Universe
- Every eligible answer: an answer collected through the API series for the snapshot's project and prompt cohort, observed inside the half-open period, whether or not an extraction covers it.
- Not measured (covered answers)
- Not measured: no covered answer tracks this entity.
- Interval
- A Wilson interval when every prompt family contributes exactly one answer to the denominator; otherwise a cluster bootstrap that resamples prompt families with replacement and recomputes the ratio of sums.
Qualified recommendation rate
- Definition version
- qualified_recommendation_rate.v1
- Unit
- A weighted proportion: a sum of commercial weights over a sum of commercial weights, from zero to one.
- Estimand
- The commercially weighted share of answers that recommend the entity as the choice or as a named alternative.
- Numerator
- The sum of commercial weights over the covered answers in which the entity is the primary recommendation or a recommended alternative.
- Denominator
- The sum of commercial weights over the covered answers that track the entity (an answer is covered when its current extraction was produced by the snapshot's extraction version set, and it tracks the entity when that extraction holds a role for it). Each weight is the prompt's cohort membership weight, a whole number from 1 to 1000.
- Universe
- Every eligible answer: an answer collected through the API series for the snapshot's project and prompt cohort, observed inside the half-open period, whether or not an extraction covers it.
- Not measured (covered answers)
- Not measured: no covered answer tracks this entity.
- Interval
- A cluster bootstrap that resamples prompt families with replacement and recomputes the ratio of sums, because answers within one prompt family are not independent.
Citation rate
- Definition version
- citation_rate.v1
- Unit
- A proportion: a count of answers over a count of answers, from zero to one.
- Estimand
- The proportion of cited answers that cite the brand's own domain.
- Numerator
- The answers with at least one citation to the controlled domain, which is the project's primary domain as recorded on the extraction.
- Denominator
- The number of the covered answers that track the entity (an answer is covered when its current extraction was produced by the snapshot's extraction version set, and it tracks the entity when that extraction holds a role for it) that carry at least one citation and record a controlled domain. Defined for the brand-role entity only.
- Universe
- Every eligible answer: an answer collected through the API series for the snapshot's project and prompt cohort, observed inside the half-open period, whether or not an extraction covers it.
- Not measured (answers with citations)
- Not measured: no covered answer that tracks this entity carries a citation.
- Not measurable (controlled domain)
- Not measurable: no controlled domain is recorded for this entity. Only the brand-role entity has one, the project's primary domain; competitor and watchlist domains are not recorded.
- Interval
- A Wilson interval when every prompt family contributes exactly one answer to the denominator; otherwise a cluster bootstrap that resamples prompt families with replacement and recomputes the ratio of sums.
Share of recommendation (weighted)
- Definition version
- share_of_recommendation.v1
- Unit
- A signed share: signed points over absolute points, from minus one to one.
- Estimand
- The entity's commercially weighted recommendation points as a share of the absolute points of every tracked entity in the same answers. It lies between minus one and one, and is negative only when discouragement outweighs recommendation.
- Numerator
- The sum, over the covered answers that track the entity (an answer is covered when its current extraction was produced by the snapshot's extraction version set, and it tracks the entity when that extraction holds a role for it), of the commercial weight times the entity's points: primary recommendation 100, recommended alternative 65, neutral shortlist 35, compared 35, contextual mention 10, discouraged or negative minus 50, not present 0 (hundredths of a point, table sor-points.v1, a product assumption).
- Denominator
- The sum, over the same answers, of the commercial weight times the absolute points of every entity the answer tracks.
- Universe
- Every eligible answer: an answer collected through the API series for the snapshot's project and prompt cohort, observed inside the half-open period, whether or not an extraction covers it.
- Point values
- sor-points.v1
- Not measured (recommendation points)
- Not measured: no tracked entity received any recommendation points in the covered answers that track this entity.
- Interval
- A cluster bootstrap that resamples prompt families with replacement and recomputes the ratio of sums, because answers within one prompt family are not independent.
Share of recommendation point values
Share of recommendation gives an entity points for the role it holds in each answer, from point table sor-points.v1:
- Primary recommendation
- 1.00
- Recommended alternative
- 0.65
- Neutral shortlist
- 0.35
- Compared
- 0.35
- Contextual mention
- 0.10
- Discouraged or negative
- minus 0.50
- Not present
- 0.00
Share of recommendation weights each recommendation role by a fixed number of points. Those points are a product assumption, version sor-points.v1, not a measurement.
A negative share means discouraged mentions outweighed recommendations in this cohort.
Uncertainty
Every measured figure carries an interval. Answers to the variants of one prompt family are asked about the same buying scenario and tend to agree, so they are not independent, and an interval that treated them as independent would be too narrow. The prompt family is therefore the unit of independence.
Mention rate and Citation rate: when every prompt family contributes exactly one answer to the denominator, the interval is the Wilson interval at a 95% confidence level. Its lower bound is rounded down and its upper bound up, and the estimate is the observed proportion, truncated toward zero.
Otherwise, and always for Qualified recommendation rate and Share of recommendation (weighted), the interval is a cluster bootstrap: 1,000 resamples, each drawing prompt families with replacement and recomputing the ratio of sums. The bounds are the 2.5 and 97.5 per cent points of the resampled values in ascending order, each read at the nearest resampled value and never interpolated. The lower bound is rounded down and the upper bound up, so a published interval is never narrower than the arithmetic allows.
The estimate is the observed value, truncated toward zero, never a resampled one. Where it falls outside the resampled bounds, the interval is widened to reach it. A prompt family that contributes no weight to the denominator is outside the resampling frame.
The resampling seed is derived from the snapshot's inputs, the metric and the entity, and is stored with the figure together with the resample count, the statistic (ratio of sums) and the resampling unit (prompt family), so every interval can be reproduced from its own stored row.
Small samples
A figure is flagged as a small sample when fewer than 30 prompt families are in its resampling frame. The flag counts prompt families, not answers, because many answers from a few families are not a large sample.
The flag sits beside the figure and never replaces it. The threshold is stored with every figure, so a later change to it never rewrites a stored figure.
Series breaks
Each snapshot records the versions and settings its figures depend on, and is compared with the newest earlier snapshot of the same project and cohort. When they differ, the snapshot records a series break under the first of these classes that applies, in this order, and names every setting that changed:
- The metric definition changed
- The extractor or its rules changed
- The project settings or prompt cohort changed
- The way observations are collected changed
A figure after a series break is not compared with a figure before it.
When figures are shown to clients
Figures reach a client only when two conditions hold. First, the extraction quality gate has passed at the exact versions that produced the figures: the extractor, the extraction prompt, the rubric, the mention rules and the citation rules. Second, every low-confidence answer with high commercial weight has been reviewed.
The gate is measured on a labelled evaluation set with at least 100 test answers. The labels in the current set were drafted and approved by AI models rather than checked answer by answer by a person. The gate passes only when every threshold below is met:
- Citation precision
- At least 0.980
- Mention precision
- At least 0.980
- Mention recall
- At least 0.950
- Recommendation macro F1, over the primary and alternative recommendation classes
- At least 0.850
Until both conditions hold, figures are for internal work only and are labelled as not validated.
This methodology states the rule and its thresholds. It does not report the status of the gate at any version.
Metrics not yet computed
Answer position, Visibility stability, Evidence coverage and Citation diversity are part of the product plan and are not computed yet.