BETAPro is not billed during beta. Lock in the price and we will honor it at launch.Freeze this price

Methodology

How we measure AI brand visibility

Every chart in BrandBanta is computed from raw scan responses, persisted as immutable rows, and aggregated transparently. We don't fabricate trendlines, hide algorithms, or average across engines without exposing the per-engine split. This page is the full definition.

How a scan works

Each scan sends a saved prompt (or an ad-hoc one) to four AI providers in parallel — ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google), and Perplexity. We route through OpenRouter for unified billing and platform parity.

Each response is persisted as an immutable row in the response table. We never modify or delete responses after the fact — every chart you see is built from the same raw rows, so re-aggregating with a future algorithm change is always possible.

Citations (URL sources the AI cited) and structured brand mentions are extracted via a regex pre-screen + LLM extraction pass, with hallucination guards on every quoted span. See "Mention detection" below.

Mention rate

Definition
Of all scans in the window, the percentage where your primary brand appeared in at least one platform's response.

Formula: scansWithAtLeastOneMention / totalScans · 100

A scan counts as "mentioning you" if any platform's response contains a verified brand_mention row for your brand. A scan that hits 4 platforms and sees you mentioned in only 1 still counts as a mention — the metric is presence at the scan level, not the response level.

Per-platform mention rate uses the same formula scoped to a single platform: responses-with-mention / total-responses, just for ChatGPT (or Claude, or Gemini, or Perplexity).

Caveat
For workspaces with very few scans (under 5 in a window), the rate is presented but flagged with a data-depth indicator. Don't draw conclusions from small samples.

Share of voice

Definition
Your brand's mention count divided by the total mention count across your brand + every tracked competitor, over a time window.

Formula: yourMentions / (yourMentions + Σ competitorMentions) · 100

Counts include every individual mention in every response, not just distinct scans. A response that mentions you twice contributes 2 to your numerator.

Caveat
Share of voice is sensitive to your competitor list. Adding a new competitor to the tracking will dilute the metric — that's mathematically correct, not a bug. If a competitor is dominant in your category, the gap is real.

Sentiment

Definition
For each individual brand mention, our LLM extractor classifies the surrounding context as positive, neutral, or negative — about your brand specifically, not the response overall.

Aggregated sentiment is the simple count of each tone over a window. We don't compute a single "sentiment score" because averaging into one number hides the texture — a 50/50 positive/negative split is not the same as 100% neutral, but a score collapses them.

Caveat
When mention detection falls back to the regex-only path (e.g. when the free-tier extractor is gated off, or the LLM call fails), sentiment is assigned as neutral with confidence 80. Look for confidence values when interpreting individual mention rows.

Mention detection

Two passes per response:

  1. Regex pre-screen. Each brand has a name + aliases + exclusion phrases. We auto-derive a hostname from the brand's website (so "nice.com" matches without the user typing it as an alias). Word boundaries on both sides — "nice job" doesn't match the brand "NICE". Case-sensitivity is configurable per brand.
  2. LLM extraction pass. Regex hits are passed as candidate spans to a structured-output call. The model is allowed to reject candidates that aren't real mentions (e.g. "nice" the adjective) and add mentions the regex missed. Each accepted mention carries a confidence score (0-100), a sentiment label, and a contextType (primary / comparison / acquisition / incidental).

Sub-brand and product names count. "NICE CXone MPower" is treated as a NICE mention, "Acme's API" as an Acme mention, "nice.com" as a NICE mention. The extractor prompt explicitly biases toward inclusion-with-lower-confidence when ambiguous — better to surface a reviewable mention than silently drop a real one.

Hallucination guard: for every extracted mention, the stored quote must exactly equal response.text.slice(spanStart, spanEnd). Mismatches are dropped silently — no fabricated quotes.

Citation tracking

Definition
When a response cites a URL source (most commonly Perplexity, occasionally Gemini, rarely ChatGPT or Claude unless explicitly searching), we persist the URL + title + a derived hostname.

Top sources cited aggregates by hostname (strippingwww.) and ranks by count. This is the "where AIs get answers in your category" view.

Sources where you're absent is the same aggregation but filtered to responses where your primary brand was NOT mentioned. This is the stronger signal — these are the domains winning answers without you.

Caveat
ChatGPT and Claude answer from training data by default and don't return citations. Gemini cites only when search-grounding is enabled. So citation data is dominated by Perplexity in practice. If you don't run scans that include Perplexity, citation metrics will be empty.

Content gap

Definition
Saved queries (prompts) where, in the last 90 days, at least one competitor was mentioned and your brand was not. The AIs are already answering these questions; they're just naming the wrong brand.

Surfaced on each own-brand detail page as the "Content gaps" card. Each row lists the query, which competitors got mentioned (with mention counts), and the last time it was scanned.

The companion metric, gap sources, applies the same logic at the citation level: which domains do AIs cite in answers where you're absent.

Caveat
Content gap is intentionally a leading indicator, not a lagging one. If a competitor was mentioned once 60 days ago and you weren't, that query shows up. The action is to evaluate whether that query represents real customer intent in your category, not to automatically write content for every one.

Discovery (untracked entities)

Definition
Every named company, product, person, or positioning phrase that the LLM extractor identifies in a response — even if it's not in your tracked brand list — is persisted to the observation table.

Discovered entities are stored per response, never deduplicated at write time. Aggregation happens at read time: group by canonical name, count by organization, rank by frequency.

When the "New competitor discovered" alert fires (≥5 observations of the same canonical name in 14 days), the entity becomes a one-click "Track this" candidate. Tracking promotes the entity into the brand table and retroactively links every historical observation as a brand mention.

Caveat
Canonical names lowercase the entity and strip common corporate suffixes (Inc, Corp, LLC, GmbH, etc.). So "Cognigy" and "Cognigy GmbH" canonicalize to the same record. This is intentional but means edge cases can collide — if "ABC Inc" and "ABC LLC" are different real companies, they'd merge.

Cost tracking

Definition
Per-scan cost in USD, computed from token usage × the model's published rate at scan time, optionally true'd-up against OpenRouter's settled charge.

Stored in micro-dollars (integer, 1 = $0.000001) on each response row to avoid floating-point drift. The dashboard "Spent" KPI aggregates this across the time window.

When a workspace uses our shared/free-tier OpenRouter key, the cost is recorded on our side and capped via the free-tier policy. When a workspace uses its own (BYOK) key, OpenRouter bills them directly and we still record the cost for visibility — but the customer's account is the authoritative ledger.

Alerts

Definition
Post-scan rules that fire push notifications when something meaningful changed. Driven by the same data the charts read — no separate "alerts data source."

Three rules in v1, each with a cooldown to prevent spam:

  • Mention rate drop:7-day rate drops by >10 points compared to the prior 7-day window (requires ≥5 scans in current). 24h cooldown.
  • Competitor surge:a tracked competitor's mention count grew >50% week-over-week with absolute ≥3 mentions. 24h cooldown per competitor.
  • New competitor discovered: an untracked brand candidate from the observation table crosses 5+ observations in 14 days. 7-day cooldown per canonical name.

Each alert delivers via the user's enabled channels (in-app, email, webhook, Slack) — opt-out per channel per alert type in your notification settings.

Data depth & cold-start honesty

LLM responses are non-deterministic — the same prompt sent twice can produce different answers. Statistical reliability comes from volume × time.

When a workspace has under 14 days of scan history, every time-series chart shows a Data depth: N days indicator. Trends drawn over short windows are honest about being preliminary.

We don't fabricate trendlines. We don't extrapolate. If you ran 3 scans yesterday and 0 the day before, the chart shows that — not a smoothed best-fit line that creates the illusion of more data than exists.

Per-engine reporting

Every metric is available per platform AND aggregated across platforms. We never average across engines without exposing the per-engine split, because ChatGPT, Claude, Gemini, and Perplexity reason differently and averaging hides that signal.

When you see a single mention-rate number on the dashboard, that's the aggregate. The platform breakdown card shows the per-platform split that the aggregate was built from.

Found something unclear, or have a question about how a specific number was computed? Get in touch. We treat methodology questions as first-class — if the answer here is wrong, we want to know.