How Arenza scores a business in AI answers: AVS and ACS.
Two scores, one probe pipeline. AVS asks whether AI surfaces you at all; ACS asks whether — once you are surfaced — the answer actually helps sell you. Every tier, every weight, and every honest limitation is documented on this page.
Generative AI assistants now mediate an estimated 30–40% of branded research queries that previously flowed through classical web search. Unlike a search engine, an assistant collapses the answer down to a single paragraph: a buyer no longer scans ten links and decides which to trust — they read one synthesized response and act. That response may or may not mention a given business, may or may not recommend it, and may or may not hand the buyer a way to act on it.
This produces a measurement gap. Brands have decades of telemetry for the search surface (rank, impression share, click-through) and roughly nothing for the AI surface. Arenza closes that gap with a probe-based pipeline that reduces every AI answer to two scores and one diagnostic pillar:
- AVS — Agentic Visibility Score. Are you seen? Whether the assistant surfaces the business, and how prominently. Whether it cites a source for what it says is reported separately, as a diagnostic — it is not part of the score.
- ACS — Agentic Conversion Score. When you are seen, does AI help sell you? How strongly the answer recommends the business and how directly it hands the buyer a next step.
- Accuracy (diagnostic pillar). Whether the assistant's stated facts about the business match a canonical reference, with severity tiers and the verbatim quote behind every flagged claim.
Credibility in this space comes from the methodology being inspectable, not from marketing copy. This page documents how the pipeline works in enough detail to be independently reproduced or critiqued.
The atomic unit, and one non-negotiable rule
Every score begins with one atomic observation:
Business × Prompt × Intent × AI Engine × Market × Language × Scan
A classifier reads prompt + answer + business and returns discrete tiers only — never a number. Application code maps tiers to points deterministically. An LLM that hands back “AVS = 73” is discarded, not rounded. This is the rule that keeps the scores auditable: every point on this page traces to a named tier, and every tier traces to something you can see in the answer text.
Both scores are calculable from the answer alone. No analytics access, no CRM, no revenue data, and no site crawl is required to produce an atomic score.
AVS — Agentic Visibility Score
AVS = Presence (40) + Prominence (60) → 0–100 business absent from the answer → AVS = 0 prompt already names the business → AVS = N/A prompt is a how-to / support question → AVS = N/A
Presence — 40 points, binary
The business is named (or unmistakably referred to) in the answer, or it is not. No subjective scoring.
Prominence — up to 60 points, structural only
Where the business sits in the answer's structure. Deliberately blind to sentiment: for “1. Acme — best overall,” prominence sees the #1 position and ignores “best overall” — that language belongs to ACS. Keeping the two apart is what stops one signal from being counted twice.
| Tier | Points | Meaning |
|---|---|---|
| PRIMARY | 60 | First or dominant business; the hero recommendation or card |
| TOP_3 | 48 | One of the first three major businesses |
| TOP_5 | 36 | Clearly featured within the first five |
| LISTED | 24 | Included meaningfully in the main answer or list |
| INCIDENTAL | 12 | Passing mention only |
Evidence support — a diagnostic, not part of AVS (since 2026-09-10)
Whether the answer backs its claims about the business with a recognizable source is still classified on every answer and still reported. It no longer moves AVS. AVS answers one question — can buyers find you — and whether the answer cites a source is a question about credibility, not findability: an answer that puts the business first and cites nothing has found it completely. Docking points there was answering a second question inside the first.
Before removing it we measured the effect across 78 brands with at least 30 judged answers each: taking evidence out of AVS moved the brand ordering by a rank correlation of 0.997. It never separated one brand from another; it only pressed every score down. So it is reported where it is actionable — as a count such as “178 of 234 answers cite no source” — and kept out of the number that says whether you are found.
| Tier | Meaning |
|---|---|
| OFFICIAL | Claims supported by an official business or product source |
| THIRD_PARTY | Claims supported by a relevant third-party source |
| NONE | No recognizable supporting source |
Boundary rule. A call-to-action link is not automatically evidence. “Book a demo → example.com/demo” is primarily a conversion-path signal and scores in ACS; it is classified as evidence support only when the answer actually uses that page to support information about the business. One link, one signal — never both.
Branded prompts — AVS does not apply
When the prompt already names the business, the assistant naming it back is guaranteed by the question rather than earned visibility. That observation is therefore AVS = N/A: it is excluded from the prompt, intent, and business AVS means and from the AVS sample size. It is neither zero nor 100. The same branded answer remains fully eligible for ACS, because how strongly the answer recommends the business and whether it provides a path to act still have commercial meaning.
Support questions — AVS does not apply
A support-shaped prompt — an owner's question whose answer is a procedure or a spec rather than a pick (“how do I connect X to a Windows laptop”, “do I need Y to use X”, warranty, returns) — is likewise AVS = N/A. A branded prompt cannot measure discoverability because the question already carries the name; a support prompt cannot measure it because the question does not ask for a business at all. Being named inside a how-to is decided by the shape of the question, not by whether buyers can find you. ACS is untouched: how the answer describes the business once it is present still has commercial consequence. Such prompts should not be in a tracked pool in the first place — the prompt-design rules forbid generating them — so this rule mostly covers questions a customer added by hand.
ACS — Agentic Conversion Score
ACS = Recommendation Strength (50) + Conversion Path (50) → 0–100 computed ONLY over answers where the business appeared business absent → ACS = N/A (never 0)
Absence is a visibility problem, and AVS already counts it as a real zero. Recording it a second time as “ACS = 0” would additionally claim the assistant discussed the business and failed to sell it — a different statement, and an untrue one. So a business the assistants never surface has a low AVS and no ACS at all.
Recommendation strength — up to 50 points
How strongly the answer helps the buyer choose this business. Identical for every business type.
| Tier | Points | Reads like |
|---|---|---|
| EXPLICIT_WINNER | 50 | “If I had to choose one, I would choose Acme.” |
| TOP_CHOICE | 40 | “Acme is the best choice for small teams.” |
| USE_CASE_RECOMMENDATION | 30 | “Acme is particularly good for small teams.” |
| POSITIVE | 20 | “Acme is a strong option.” |
| CONSIDER | 10 | “Acme is another option worth considering.” |
| NEUTRAL | 0 | Purely descriptive — “Acme is a project-management tool.” |
A negative recommendation scores 0 and is additionally stored as a sentiment diagnostic — it never becomes a negative number.
Conversion path — up to 50 points, commerce and SaaS scored separately
How directly the answer lets the buyer act. The tiers and points are the same for every business, but what counts as a conversion signal differs by business type — an e-commerce store converts when an agent can reach a purchasable product, while a SaaS product converts when an agent can reach a sign-up, trial, or demo. Scoring both against one vocabulary would systematically misread one of them, so the classifier applies the vocabulary that matches the business:
| Tier | Points | Commerce signal | SaaS signal |
|---|---|---|---|
| DIRECT_ACTION | 50 | The buyer or agent can buy now: checkout, add-to-cart, a purchasable product card with price and stock | The buyer or agent can start now: sign-up, account creation, or a booking that completes |
| CTA_WITH_OFFICIAL_LINK | 40 | “Buy now →” / “Shop →” with a clickable official product link | “Start free trial →” / “Book a demo →” with a clickable official link |
| CTA_NO_DIRECT_PATH | 30 | A buy suggestion with no usable path — named retailer, no link; “check availability” | “Contact sales” / “request a quote” with no usable path |
| COMMERCIAL_DESTINATION | 20 | Direct access to a product or price page | Direct access to a pricing or plans page |
| GENERIC_NEXT_STEP | 10 | “Check it out” / “visit Acme” | “Learn more” / “visit Acme” |
| NONE | 0 | No next step | No next step |
The strongest conversion path present in the answer is the one that scores.
The coverage rule
Because ACS exists only where the business appears, Arenza never displays an ACS without its coverage:
ACS: 82 · present in 24 of 50 answers · coverage 48%
An ACS of 95 across 3% of answers and an ACS of 95 across 90% of answers describe opposite businesses. AVS captures that difference; the coverage line explains it. Any surface that prints an ACS alone is treated as a bug.
Worked example
Prompt "What's the best tool for X?"
Answer "1. Acme — Best overall.
If X is your main workflow, Acme is one of the strongest choices.
Start Free Trial →"
Presence PRESENT 40
Prominence PRIMARY 60
AVS = 100
Recommendation TOP_CHOICE 40
Conversion path CTA_WITH_OFFICIAL_LINK 40
ACS = 80
Evidence support NONE diagnostic — reported, not scoredNote where the trial link landed: it raised ACS, not AVS. The answer puts the business first, so it is fully found; it recommends the business strongly and hands the buyer a next step — but cites no source, which is reported as an evidence diagnostic rather than docked from visibility. Two scores, two different problems to fix; one diagnostic that says where the credibility gap is.
How atomic scores roll up
Observations aggregate in three steps, and each step refuses to let a countdecide importance:
| Level | Rule | Because |
|---|---|---|
| observation → prompt | equal weight per observation | more engines and markets sharpen the prompt's estimate |
| prompt → intent | equal weight per prompt | scanning one prompt more often must not silently promote it |
| intent → business | equal weight per intent | a prompt is a sample of a buyer intent, not an item on a checklist — an intent with 13 phrasings must not out-vote one with 3 |
- Absent observations stay AVS = 0, so visibility frequency is already inside the AVS average — it is never multiplied by coverage a second time.
- ACS averages only present, classified answers, and travels with its coverage at every level.
- Search volume never weights a score. It decides which questions we track — never what a measurement is worth. Volume coverage is patchy and drifts on its own; a score whose weights depend on how much volume data happened to be available is not comparable across brands or across weeks.
- Every slice uses the same arithmetic. Per-engine, per-market, and per-intent scores are the same roll-up applied to a filtered set of observations. There is no engine-specific or market-specific formula.
The two scores also compose into a diagnosis. High AVS with low ACS means the assistants see the business but do not sell it — a positioning, proof and conversion-path problem. Low AVS with high ACS means the answers that do surface it sell it well, but it is not surfaced enough — a coverage, content and authority problem. The product always connects a score gap to an action.
Sampling design
For each tracked brand we populate a coverage grid of (assistant × market) cells. Three assistants are in scope today: ChatGPT (OpenAI), Gemini (Google), and Perplexity. Four markets are in scope: United States (en-US), United Kingdom (en-GB), Germany (de-DE), and Japan (ja-JP). That yields 3 × 4 = 12 cells per brand. Markets are realized through prompt-language localization plus, where the assistant API supports it, a system-locale hint.
Within each cell the pipeline draws between 50 and 200 samples, scaled upward for high-volume brands and downward for cold-start brands where cost dominates marginal precision. The standard normal-approximation interval at the worst-case probability p = 0.5 yields N ≈ 96 for a ±10% margin of error and N ≈ 384 for ±5%; the per-cell sample size is recorded with every result so confidence intervals can be reconstructed downstream rather than asserted by us in marketing copy.
Across a quarterly research cohort of 1,000 brands at the standard probe count, this pipeline produces approximately 24,000 (brand, assistant, market) data points per quarter, which is what the public benchmark reports are built on top of.
Prompt generation
Probes are constructed to mirror the questions a buyer asks during genuine research, not the questions a brand wishes to be asked. Two design rules govern every prompt set:
- Unbranded majority. At least 70% of probes per brand do not contain the brand's name, the brand's product names, or any brand-specific jargon. Branded probes return tautological positive responses (the brand is the subject of the question, so it appears in the answer) and therefore measure little about organic surfacing. The 70% floor is enforced programmatically by the prompt-writer agent that drafts the matrix; the unbranded share is logged per brand for audit.
- Buyer voice. Probes are written in the first person from the perspective of a buyer with a stated job-to-be-done, not in the third person from an analyst's perspective. “I'm setting up a small video studio and need…” produces measurably different retrieval and different mentions than “list the top brands for video studios.”
Inside each cell, prompts span three orthogonal axes: 5 personas (role, seniority, organization size, budget tier, technical depth, derived from the brand's stated ICP), 5 intent classes (discovery, comparison, validation, problem-solving, purchase), and 3 temperature samples (or three independent draws separated in time, where the assistant API does not expose a temperature control). The Cartesian product yields 5 × 5 × 3 = 75 prompt variants per (brand, market) pair, each issued to each assistant in scope. Near-duplicate phrasings are collapsed by a semantic dedup pass before the matrix locks, so the intent-level averages above are not skewed by one question asked seven ways.
Mention extraction
Every returned answer is parsed for brand mentions in two stages. The result feeds the Presence and Prominence components of AVS.
The first stage is a deterministic regex layer keyed on the brand name and the brand's known aliases — the legal name, common abbreviations, the domain stem, and product family names. This stage is fast, cheap, and recovers the large majority of mentions across the multi-LLM coverage grid.
The second stage is an LLM-judge fallback that runs only when the regex layer is ambiguous (e.g., the brand's name is a common English word like “Apple” or “Square,” or the answer refers to the brand by a paraphrase like “the German camera company that makes the SL2”). The judge is given the answer, the brand's canonical description, and a strict yes/no/uncertain rubric. Judge calls are themselves sampled at low temperature and aggregated by majority. The split between regex-resolved and judge-resolved mentions is recorded per cell so the pipeline's reliance on the more expensive, more drift-prone judge layer is auditable.
Accuracy pillar: wrong-claim detection (3 severity tiers)
Alongside the two scores, a separate accuracy pillar checks what the assistants actually say. For every mention, the surrounding sentence is parsed for assertions about the brand: founding year, headquarters, product features, pricing, customer base, and other structured facts. Each parsed assertion is compared against a canonical reference for the brand, constructed in priority order from (1) the brand's own llms.txt / llms-full.txt if published, (2) a structured profile maintained by the brand or its agency inside the Arenza dashboard, and (3) public structured-data sources such as Wikidata.
Discrepancies are recorded as wrong claims and assigned one of three severity levels:
- Critical (factual error) — a verifiable fact stated incorrectly: wrong founding year, wrong headquarters city, wrong price, wrong product capability. Acted on first because it directly misleads a buyer about a hard fact.
- High (positioning error) — a fact stated in a way that materially mischaracterizes the brand's position: a premium offering described as “budget,” a B2B-only product described as consumer, a flagship described as deprecated.
- Medium (attribution error) — a fact correctly stated but attributed to the wrong entity: a feature of a competitor attributed to the brand, or vice versa. Important to fix because it dilutes brand-attribution share even when individual statements are technically true.
Every wrong-claim record includes a verbatim quote of the offending sentence and a pointer to the canonical reference contradicting it, so brand-side review is auditable rather than opaque.
Diagnostics: share of voice and friends
A family of supporting metrics — mention rate, average position, share of voice, top-3 rate, evidence support (official / third-party / none), CTA rate, engine and market consistency — explains why an AVS or ACS is high or low. They are deliberately not part of either formula, and are never labelled AVS or ACS on any surface: a mention rate tells you whether you were named; AVS tells you whether being named was worth anything.
The most requested diagnostic is share of voice. For a brand b, assistant a, market m, it is the intent-weighted mean mention rate across the prompt matrix:
SoV(b, a, m) = Σ_i w_i · ( 1 / |P_i| ) · Σ_{p ∈ P_i} mentions(b, p, a, m) / N_{p,a,m}
where:
i ∈ {discovery, comparison, validation, problem-solving, purchase}
w_i per-intent weight (uniform by default; configurable per brand)
P_i set of prompts in intent class i
N_{p,a,m} per-cell sample size for prompt p in (assistant a, market m)Competitive share-of-voice replaces the numerator with mentions(b, p, a, m) over the sum of mentions across the brand and a configured competitor set, so “our 18% vs their 22%” comparisons are computed against the same prompt matrix and the same sample sizes.
Cross-LLM divergence index
A central observation motivating the multi-LLM design is that the assistants disagree, sometimes substantially, on which brands surface for a given prompt. Reporting on only one assistant systematically under-reports cross-assistant variance and conceals divergence the brand needs to act upon.
We quantify the disagreement with a cross-LLM divergence index based on the Jensen-Shannon divergence between assistant answer distributions. For a fixed prompt p in market m, let Da be the empirical distribution over the brands mentioned by assistant a across the cell's samples (treating “no brand mentioned” as a distinct outcome). For each pair of assistants we compute:
JSD(D_a, D_a') = ½ · KL(D_a || M) + ½ · KL(D_a' || M)
M = ½ · (D_a + D_a')
CLDI(p, m) = ( 1 / 3 ) · Σ_{a < a'} JSD(D_a, D_a') over all C(3,2) = 3 pairsThe Jensen-Shannon divergence is symmetric, bounded in [0, log 2], and well-defined even for distributions with disjoint support. CLDI close to 0 indicates the assistants surface the same brands at the same rates; CLDI close to log 2 ≈ 0.693 indicates they surface entirely disjoint brand sets. A brand-level CLDI is the prompt-weighted mean over the prompt matrix; a high CLDI brand has a fragmented assistant footprint and is structurally exposed to single-assistant blind spots.
Limitations and transparency
Honest limits of the methodology:
- Probe-based, not census. The pipeline measures a structured sample of (prompt, assistant, market) cells, not the population of every answer ever generated. The roll-up estimators are unbiased under the design assumptions, and per-cell N is recorded so confidence intervals can be reconstructed.
- ACS depends on classification coverage. The recommendation and conversion tiers require the classifier to have read the answer. Where only part of a brand's visible answers have been classified, the ACS describes that subset, and the classified share is reported alongside it. Unclassified answers are shown as unmeasured — never scored as zero.
- Stochasticity is mitigated, not eliminated. Repeating the same probe yields different answers. Per-cell sampling tightens the within-cell estimator, but a single user's single query at a single moment may still observe an outcome away from the cell mean. The methodology measures a distribution, not a guaranteed individual experience.
- Model updates break temporal comparability. An assistant's underlying model may change between observation windows, breaking direct period-over-period comparison. We record the assistant identifier and observation timestamp; we do not yet have a principled framework for normalizing measurements across model updates. This is honestly an open problem.
- Locale via prompt language plus system locale. Market targeting is realized through prompt-language localization and, where the API supports it, an explicit system-locale hint. This does not fully capture per-user personalized assistant answers driven by sign-in state, location, and prior conversation history.
- Wrong-claim detection requires canonical truth. Accuracy measurement compares assistant assertions to a brand's canonical reference. Brands without a published
llms.txt, without a maintained dashboard profile, and without a Wikidata presence are reported as a measurement gap rather than imputed. - Sampling biases. The persona / intent matrix is derived from the brand's stated ICP; if the ICP description is wrong, the matrix systematically misweights the buyer-funnel coverage. We mitigate this by surfacing the matrix to the agency for review before the first scan locks.
Collection windows. Quarterly research cohort scans run on a published calendar; ad-hoc per-brand scans run continuously. The dashboard always shows the scan timestamp alongside the score so a number is never read out of context.
LLM versions. The pipeline records the assistant identifier and, where exposed by the API, the underlying model snapshot. Comparing scores across observation windows that span a model update should be done with that caveat in mind; the dashboard surfaces a model-change marker on relevant trend charts.