00Full measurement methodology

How Lokrix AI measures AI visibility

Enter a URL. Lokrix AI auto-generates a prompt universe from real buyer questions, applies a different prompt strategy to each of 7 answer engines based on how each engine actually retrieves — Bing rank, Google organic, live web, or parametric knowledge — then samples every engine K times across temperature, persona, locale, and device using a sequential Wilson stopping rule. Each run produces a rich per-draw signal set. Across K runs, Lokrix AI reports a calibrated Monte-Carlo presence probability with a 95% confidence interval, 50+ derived metrics, and SHAP-grounded attribution — never a single-run number, never an invented score.

01URL → CrawlBrand + industry
02Prompt universeIntent fan-out
03Engine tuning7 playbooks
04K × Axis samplingSequential Wilson
05DetectionURL · NER · semantic
06Metric catalog50+ signals
07Score + 95% CICalibrated probability

01Build the prompt universe

Real buyer questions, auto-generated from your URL — not keyword phrases.

Lokrix AI crawls your URL, extracts brand, product, and industry context, then runs a multi-technique prompt generator that produces intent-complete natural-language prompts — the actual questions buyers type into ChatGPT, Perplexity, and Google AI Mode, not cleaned-up keyword phrases.

Each prompt is tagged by intent class: informational, commercial-investigation, transactional, comparison, or trust-safety. The universe is built by intent laddering — awareness through consideration to decision — and de-duplicated on embedding similarity (not just string matching), so paraphrases collapse into one representative prompt.

The prompt universe is editable and persistent: add, remove, or weight prompts by funnel stage, and Lokrix AI tracks them all over time and across every cadence. Fan-out sub-queries captured from Google AI Mode and Gemini are harvested and fed back as new universe seeds at each cadence.

Auto-generated prompt examplesillustrative
  1. 01
    Best enterprise analytics platform for data teams?commercial-investigation
  2. 02
    Which project management tools do engineering leads prefer?informational
  3. 03
    Compare [Brand] vs alternatives for B2B growth teamscomparison
  4. 04
    What security framework do analysts recommend for SaaS companies?trust-safety
  5. 05
    How to get started with [Category]?transactional

Real prompt universes derive from your brand, industry, and competitor set. Intent-class tags drive per-engine prompt tuning in Step 03.

02Per-engine prompt engineering

Seven engines. Four retrieval classes. A different strategy for each — because what ranks on Bing is not what ranks on Google.

Lokrix AI does not submit the same prompt text to all engines. Each engine grounds on a different index — Bing, Google, live web, X, or parametric knowledge — with different weighting functions. The prompt strategy, the sampling axes, and the priority metrics are tuned per engine to extract the maximum signal from what that engine’s retrieval substrate actually responds to.

Engine classes: Bing-grounded (ChatGPT), Google-grounded (AI Overviews, AI Mode, Gemini), Live-search / social (Perplexity, Grok), and Model-as-ranker (Claude) — one extraction spine, four per-class playbooks.

ChatGPT

OAI-SearchBot · GPTBot

Bing-grounded

Retrieval: Bing index — 87% of citations match a Bing top-10 result

Primary levers: Bing rank and freshness are the primary citation levers

Prompt techniques

  • Intent fan-out with Bing-realistic phrasings across 5 intent classes: informational, commercial-investigation, transactional, comparison, trust-safety
  • Locale ladder — Bing is market-partitioned; a brand can be surfaced en-US and invisible de-DE
  • Copilot grounding-query mining — Copilot exposes its actual Bing grounding queries, which Lokrix feeds back as ChatGPT variants to close the loop on ChatGPT's opaque retrieval
  • Citation-forcing variants kept in a separate DIAGNOSTIC stratum — never blended into the organic headline probability

Priority metrics

  • Bing rank of the cited URL (shadow-SERP anchored to Copilot grounding queries)
  • Citation-rate vs mention-rate gap ("known but not Bing-surfaced" is a different problem from "not known at all")
  • Freshness of cited pages — Bing rewards recent content

Google AI Overviews

Google-Extended · Googlebot

Google-grounded

Retrieval: Google index — 54% of citations in Google organic top-20; AIO fires selectively

Primary levers: AIO does not fire for every query — render-rate is a first-class metric. Google's retrievalContext already exposes the shadow organic rank.

Prompt techniques

  • Entity-centric prompts — AIO favors topically-authoritative entities over exact-match keywords
  • PAA (People Also Ask) and related-questions seeding to mirror the AIO trigger pattern
  • Locale and device variation — AIO personalizes heavily by geography and device type
  • Comparison and recommendation elicitation where AIO and AI-Mode diverge most

Priority metrics

  • Render-rate — fraction of draws where AIO fired at all (P(surfaced | rendered) vs unconditional P(surfaced))
  • Google shadow organic rank of cited URLs (captured from retrievalContext.ranking in the AIO adapter)
  • AIO↔AI-Mode divergence — per-(prompt, locale, device) Jaccard and target quadrant (in-both / AIO-only / AI-Mode-only / neither)

Google AI Mode

Google-Extended · Googlebot

Google-grounded

Retrieval: Google index — same base as AIO, different selection head; ~14% URL overlap with AIO

Primary levers: Spawns fan-out sub-queries per draw, decomposing intent into sub-questions (captured as groundingQueries)

Prompt techniques

  • Fan-out query harvesting: seed prompt → realised fan-outs → persist as new prompt-universe seeds for the next cadence
  • Comparison and recommendation elicitation to surface maximum divergence from AIO
  • Device and locale variation — AI Mode behaviour differs meaningfully across contexts

Priority metrics

  • Fan-out count and decomposition ("Google decomposed into N sub-questions; target appears in K of N")
  • AIO↔AI-Mode target quadrant — the structural divergence between the two Google surfaces
  • Shared Google shadow SERP rank (one search call per cell, shared with AIO and Gemini)

Perplexity

PerplexityBot

Live-search / social

Retrieval: Live Google + Bing blend; reads cited pages; ~47% Reddit/UGC by citation share

Primary levers: Freshness-weighted. Sonar returns the richest sensor payload: per-result date, snippet, and a clean ordered citation list vs. a retrieved-but-not-cited set.

Prompt techniques

  • Freshness-forcing paired variants (neutral vs "as of 2026 / latest") — measures recency-elasticity
  • Reddit/forum-intent phrasings to probe UGC-dependence and identify which sub-communities Perplexity draws from
  • Follow-up "cite your sources" 2nd turn — promotes the target from retrieved to cited under scrutiny
  • Source-count elicitation to widen the candidate pool and reveal retrieval vs selection gaps

Priority metrics

  • Numbered citation rate + retrieved-not-cited rate — the single most diagnostic native metric (retrieved #7 but not cited → "extracted not selected")
  • Citation-domain composition with UGC/authoritative/owned/competitor split
  • Freshness — median citation age from the date field on each search result

Gemini

Google-Extended · Googlebot

Google-grounded

Retrieval: Google Search grounding; schema-sensitive; entity-clarity-weighted

Primary levers: Entity clarity — how well the brand is defined in the knowledge graph — is the primary lever alongside Google organic rank

Prompt techniques

  • Entity-centric prompts and entity-disambiguation variants
  • Structured-output JSON elicitation to capture recommendation rank and rationale
  • Schema-rich variants to test structured vs prose answer modes for citation behaviour

Priority metrics

  • Google shadow rank (shared SERP call across AIO, AI Mode, and Gemini)
  • Entity-mention rate — brand named in the answer without a cited URL (a knowledge-graph signal)
  • Grounding query capture — the actual search queries Gemini ran to formulate the answer

Claude

ClaudeBot

Model-as-ranker

Retrieval: Parametric knowledge + shallow live search via ClaudeBot

Primary levers: Factual density and entity clarity are the primary levers. Claude hedges explicitly and explains its own selection criteria on request — the model is its own attribution dataset.

Prompt techniques

  • Ranked-recommendation elicitation: structured JSON [{name, url, rank, reason, confidence}] — forces an ordered list even when the brand is named-but-not-cited
  • CoT elicitation — surfaces the selection criteria the model used as feature-level attribution
  • Self-critique: "what would make you more confident recommending X?" — the engine states its own gap to close
  • Self-consistency over K draws: modal recommendation list + agreement rate as a stability metric

Priority metrics

  • Recommendation rate + rank-in-list (the primary KPI for model-as-ranker engines)
  • Stated reasoning mapped to the shared lever taxonomy and reconciled against SHAP
  • Hedge rate + self-reported confidence per draw

Grok (xAI)

Beta

X / live web

Live-search / social

Retrieval: X (Twitter) posts + live web; freshness-weighted; social-corroboration-heavy. Beta — Grok is live via a web-search fallback, so its draws are kept conservative and not yet counted in headline scores.

Primary levers: Social proof and X engagement are the primary levers; over-indexes recency and social corroboration

Prompt techniques

  • X/social-share phrasings that probe which brands are being talked about on X
  • Recency-forcing variants to measure freshness-elasticity
  • Comparison prompts: win/loss + stated reason per rival

Priority metrics

  • X/social-source share of cited domains
  • Recommendation rate + rank-in-list
  • Freshness of cited pages — recency is weighted more heavily than on index-grounded engines
Monte-Carlo sample cluster (illustrative). Each dot is one engine run. The cluster density represents a 62% presence probability with a 95% confidence interval of [54%–70%]. Not a live measurement.
62%presence probability · 95% CI [54%–70%]illustrative · not a live measurement

03Monte-Carlo sampling

A single AI answer is a sample from a distribution. Lokrix AI collapses the distribution into a probability — with a CI.

For every (prompt × engine) cell, Lokrix AI runs K draws with browsing and grounding enabled, varying across four axes: temperature ladder (0.2 / 0.6 / 0.9), buyer persona, locale, and device type. Temperature spread across the ladder is reported as a volatility signal — a brand that only appears at low temperature has fragile, low-confidence presence.

The number of draws is controlled by a sequential Wilson stopping rule: while the 95% CI width exceeds 2ε (default ε=0.10), take another draw. Minimum K=8; maximum K=60 (K=40 for token-heavy reasoning engines). Cells with decisive outcomes — presence probability above 0.8 or below 0.2 — exit early. This concentrates cost exactly where uncertainty is high.

Across K draws, the raw mention counts become a Monte-Carlo presence probability — the fraction of draws where the target was surfaced. The cluster diagram on the left illustrates what this looks like: many faint draws (each one an engine run), a dense core at the point estimate, and a dashed ring at the 95% CI boundary.

Provenance on every draw: live, cached, or predicted. A sampled result is never presented as a live measurement. A single run is never presented as a probability.

04Mention detection

Three detection layers. Five surface classes. A refusal is not a non-mention.

On every draw, Lokrix AI runs three parallel detection layers over the answer text and citation cards. Each layer produces its own evidence; the highest-confidence detection determines the draw’s surface class.

01
URL / domain match
Exact URL or domain substring found in the engine's citation cards. Zero false positives by design — the hardest evidence layer.
02
Brand NER
Named entity recognition over the answer text. Fuzzy-matched (Levenshtein edit distance ≤ 1) with token-overlap scoring (θ=0.6). Covers the brand name, product name, and common aliases.
03
Semantic similarity backstop
TF-IDF relevance score between answer passages and the brand's descriptor tokens. Catches paraphrase mentions — when the engine describes the brand without naming it. Current outputs carry a lower confidence weight; a cross-encoder upgrade is on the roadmap.

Refusals are excluded from the presence denominator — an engine that refuses to answer the prompt is not a "no-mention." Refusal rate is tracked separately as K4_refusalRate and flagged in quality metadata, so the presence probability is never biased upward by silently dropping refusals from K.

A periodic negative-control battery (fabricated brand names) measures the detector’s false-positive rate across the URL, NER, and semantic layers independently. Reported presence is bias-corrected by subtracting the measured FPR from the raw rate.

Surface classes — strongest to weakest evidence
  1. Cited

    URL in a citation card. The engine explicitly retrieved and cited your page. Hardest evidence.

  2. Linked

    Inline hyperlink to your domain in the answer text. Strong retrieval evidence even without a citation card.

  3. Recommended

    Recommendation language — "use X", "best for Y", "I suggest X". The buying-intent surface. What brands most want.

  4. Named

    Brand named in prose — the awareness floor. High name-rate with low citation-rate = "known but not indexed for this query."

  5. Semantic

    Semantic similarity only — the engine describes your category and use case without naming you. A latent presence signal.

The citation-rate vs mention-rate gap is diagnostic: high named + low cited = "known but not indexed for this query" — a different lever problem from "not known at all."

05The metric catalog

50+ calibrated metrics organized into families — most extracted from signals the engine already returns.

From every draw, the adapters capture a rich dict — grounding queries, citation metadata, answer shape, rendering status, source snippets with dates. Most of the 50+ metrics in the Lokrix AI catalog require no additional API calls: they plumb through data that engines already return and the naive pipeline would discard. Metrics that require a new call (page fetch, LLM-judge, embedding) are budget-clamped and metered against the tenant’s credit balance.

Every rate in every family carries a Wilson 95% CI. Every aggregate across draws is Beta(2,5)-shrunk before combining to prevent 0/1 overconfidence at low K. Every mix (surface-type histogram, provenance breakdown) reports counts + n, never a lone percentage that hides bimodality.

Family A

Presence / probability

The measurement spine. Every rate carries a 95% CI and a Beta-shrunk aggregate.

  • p̂ — raw presence probability (successes / K)
  • p̃ — Beta(2,5)-shrunk posterior mean; prevents 0/1 overconfidence at low K
  • Wilson 95% CI [lo%, hi%] on every single rate
  • Presence variance — propagated into the headline score CI
  • Sample-adequacy flag — "low_sample" below K=10, plus K needed to reach ε=0.10
Family B

Surface-type rates

What kind of surfacing happened — not just whether. Each rate has its own CI.

  • Citation rate — URL in a citation card (highest value)
  • Link rate — inline hyperlink in the answer text
  • Recommend rate — recommendation language in the answer (0.85 confidence label)
  • Name rate — brand named in prose (awareness floor)
  • Strong-class rate — fraction of surfaced draws that were cited or linked
Family C

Position / rank

Where in the answer the brand appeared. Earlier positions carry disproportionate weight.

  • First-citation ordinal — rank of first appearance
  • PAWC — position-adjusted weight: 1/log2(rank+1)
  • Average ordinal with bootstrap CI
  • Top-citation rate — fraction of runs where the brand was cited first
  • Rendered rate (AIO/AI-Mode) — fraction of draws where the surface fired at all
Family D + E

Share of voice + competitor delta

Presence relative to the field, computed per draw so the CI is honest.

  • Presence-ratio SoV — p̃_target / (p̃_target + Σp̃_competitors)
  • Position-weighted SoV — PAWC-adjusted so #1 and #5 are not treated equally
  • Head-to-head win rate — paired draw-level comparison; McNemar test for significance
  • Exclusive-mention rate — draws where the brand appears and no competitor does
  • Competitor displacement opportunity — beatable dominant rivals from trend + delta
Family F

Sentiment / framing

How the brand is described when it appears. Labeled a lexical stand-in; an LLM-judge upgrade path is documented.

  • Sentiment polarity — lexical classifier over the brand mention window, mean per draw
  • Sentiment mix — positive/neutral/negative histogram with counts
  • Framing keywords — adjectives and roles around brand spans (evidence, not a score)
Family G

Cited-domain provenance

Who gets cited alongside or instead of the target — the corroboration graph.

  • Source provenance mix: owned / earned / competitor / UGC / reference
  • UGC dependency (Reddit/forum share — critical for Perplexity, ~47%)
  • Reference corroboration rate — Wikipedia/G2 co-cited with the target
  • Shadow-index overlap — ranks high in Google/Bing but is not cited (extractability gap)
Family I + M

Freshness + volatility

How recent the cited sources are, and how stable the answer is across draws.

  • Freshness — median age of cited pages (critical for Perplexity and ChatGPT)
  • Recency bias per engine — correlation between citation rank and page age, validated empirically
  • Presence consistency — 1 − 4·p̂·(1−p̂); a coin-flip at p̂=0.5 vs a reliable cite
  • Cited-set Jaccard stability — mean pairwise overlap of cited-domain sets across draws
Family K

Answer-shape

The structure of the engine's answer itself — refusals, length, citation density.

  • Refusal rate — excluded from the presence denominator (a refusal is not a "no mention")
  • Answer length and citations-per-100-words density
  • Fan-out / grounding-query count (new prompt-universe seeds)
  • Answer format type — list vs prose vs table (list answers surface more brands)
Family N

Trend / momentum

Presence over time. Regression alerts fire on real losses, not noise.

  • Presence trend slope (per engine, per prompt)
  • Momentum — EMA short − long delta
  • Regression alert — two-proportion test or CI non-overlap → Slack/email/webhook
  • Competitor momentum — are rivals catching up or falling behind?

06Per-AI reasoning attribution

The “why” behind every score — grounded in each engine’s specific retrieval substrate, not generic advice.

Attribution runs through two complementary channels, reconciled into one grounded explanation per score.

Channel A — Elicited (ask the engine)

Chain-of-thought elicitation, self-critique ("what would make you more confident recommending X?"), structured JSON with rationale, and follow-up source-probing on decisive draws. For Bing-grounded engines, the elicited self-report is verified against Copilot’s actual grounding queries to flag confabulation. These draws are gated by cost and marked conversational — never independent draws in the Monte-Carlo.

Channel B — Inferred (compute, don’t ask)

Computed from already-captured signals: Bing/Google shadow rank (the strongest single driver), freshness, owned-vs-earned corroboration count, position-in-answer, and competitor crowding. A linear SHAP surrogate over these features produces per-feature contributions; TreeSHAP over the trained gradient-boosted model is the upgrade path.

Stated attribution (what the engine says it considered) is compared against the SHAP-computed drivers. Agreement → high confidence. Divergence → surfaced with a note and used as a calibration input for per-engine model weights. Over K draws, stated weights are aggregated as stability-weighted attribution share(draws citing lever) × mean(confidence) — so a single-draw anecdote never drives a recommendation.

Counterfactuals are credible because the levers are rank-driven: "Move Bing rank #12→#5: predicted +Δ presence." Perplexity gap-decomposition: "Target retrieved at position #7 but not cited; the two cited rivals each carried a statistic and were 60 days fresher." Every predicted uplift is labelled simulated. llms.txt is never surfaced as a lever — no measured effect exists in current research.

Shared lever taxonomy

The vocabulary used by stated attribution, SHAP-computed drivers, and optimization recommendations — all three speak the same taxonomy so they are directly comparable.

Factual density / statistics
Concrete numbers with sources; Princeton +40% citation lift
Direct quotations
Attributed expert quotes; Princeton +41% citation lift
Cited sources
Links to primary sources in the content; Princeton +30%
Third-party corroboration
Reddit/Wikipedia/G2 co-citations; top driver across all engines
Freshness / recency
Page publish/update date; dominant for Perplexity and ChatGPT
Entity clarity
How unambiguously the brand is defined in structured data; critical for Gemini and Claude
Official / authoritative domain
Domain authority and trust signals; Bing-grounded engines weight heavily
Topical relevance
Content alignment to the prompt intent; measured by cross-encoder relevance
Brand reputation
Aggregate signal from earned-media mentions, reviews, and social proof
Pricing / feature fit
Explicit feature comparison coverage; important in commercial-investigation prompts
Social proof (X for Grok)
X post engagement and social corroboration; the primary Grok lever

Source: Princeton GEO study (KDD 2024) + Answer-engine citation-source analyses (2024–2026), summarised in CLAUDE.md §4.

07Statistical honesty

Every number is a calibrated probability with a 95% CI. Non-determinism is a first-class measurement, not a bug.

These are the non-negotiable protocols that make Lokrix AI results credible rather than convenient:

  1. 1

    Wilson 95% CI on every rate

    wilson_interval(successes, K, z=1.96) on every single metric — citation rate, link rate, recommend rate, SoV, position ordinal. Never a bare point estimate displayed without its interval.

  2. 2

    Beta(2,5) shrinkage before aggregating

    Every presence aggregate uses (successes+2)/(K+7) — a Beta(2,5) posterior mean — instead of raw p̂, preventing 0/1 overconfidence at low K. A "low_sample" flag appears below K=10.

  3. 3

    Sequential Wilson stopping — cost-honest sampling

    Samples continue while CI width > 2ε (ε=0.10, adjustable). Decisive cells (p̂ > 0.8 or p̂ < 0.2) exit early. K_min=8, K_max=60 (40 for token-heavy reasoning engines). This concentrates cost where it matters without over-sampling decisive outcomes.

  4. 4

    Negative controls — measured FPR, bias-corrected presence

    A periodic battery of fabricated brand names measures the detector false-positive rate across the URL, NER, and semantic layers independently. The reported presence rate is bias-corrected: FPR is subtracted and the CI is widened accordingly. Specificity is always measured alongside sensitivity.

  5. 5

    Refusals excluded from the denominator

    An engine that refuses to answer is not a "no-mention." Refusals are excluded from K in the presence denominator and reported separately as refusal_rate. Silently counting refusals as non-mentions would bias presence upward.

  6. 6

    Mode separation — organic only in the headline

    Only ORGANIC draws (recommendation-elicitation, intent fan-out with no brand in the prompt) enter the headline composite. DIAGNOSTIC-ELICITED, CONVERSATIONAL, CALIBRATION, and ROBUSTNESS draws are reported in separate breakdowns. Blending a draw where the brand was named in the prompt into the organic probability is the cardinal violation.

  7. 7

    Provenance on everything

    "live" / "cached" / "predicted" badge on every data point. Live probabilities carry an "uncalibrated" flag until the Platt/Isotonic calibration map has been fit from historical predicted→realized outcomes. Mock and illustrative data is always labelled and never dressed up as a live measurement.

  8. 8

    Calibration: Brier, ECE, AUC

    The predictive scoring model is evaluated on Brier score (mean squared error of probabilities), Expected Calibration Error, AUC, and a reliability curve. Platt scaling is the default; isotonic regression is used only with ≥200 samples and a ≥2e-2 improvement in CV-ECE. The goal is not the sharpest model but the most honest one.

illustrative

The shaded band is the honest 95% CI. It narrows as K increases and as the true probability moves away from 0.5 toward decisiveness.

Presence probability over time

With its 95% CI band

Not a live measurement — shows the shape of a real Lokrix AI report.

Engine grounding — real citation-source data

Where citations come from, by engine.

ChatGPT
87%BingBing top-10 rank + freshness
Google AI Overviews
54%GoogleGoogle top-20; separate from AI Mode
Google AI Mode
GoogleShares a cited URL with AIO only ~14% of the time
Perplexity
47%Google + Bing (live)Over-weights Reddit (~47% of top citations) + freshness
Gemini
GoogleGoogle Search grounding; schema-sensitive
Grok (xAI)
X + live webGrounds on X posts + live web; freshness-weighted (beta — live via a web-search fallback; draws not yet counted in headline scores)
Claude
Parametric + ClaudeBotFactual density and entity clarity are the primary levers; parametric knowledge with shallow ClaudeBot live search

Source: Answer-engine citation-source analyses (2024–2026), summarised in CLAUDE.md §4. Figures where published; null where the engine architecture is documented but no single headline percentage has been published.

Glossary

Key terms — real definitions for AI extraction and transparency.

GEO — Generative Engine Optimization
The practice of making a brand more likely to be cited, linked, recommended, or named in AI-generated answers — ChatGPT, Google AI Overviews, AI Mode, Perplexity, Gemini, Grok, Claude. GEO is to answer engines what SEO is to the ten blue links.
Presence probability
A calibrated estimate of the fraction of engine runs in which a given URL, brand, or product is surfaced for a given prompt. Expressed as a percentage and always reported with a 95% confidence interval. Never derived from a single run.
Confidence interval (95% CI)
The range [lo%, hi%] within which the true underlying probability is expected to fall in 95 out of 100 repetitions of the sampling procedure. A wide CI means more uncertainty; a narrow CI means more evidence. Always displayed alongside the point estimate — never hidden.
Surface class
The five levels of how a brand appears in an AI answer: cited (URL in a citation card), linked (inline hyperlink), recommended (recommendation language), named (prose mention only), and semantic (descriptor match without an explicit name). Each class has its own rate and CI.
Share of voice (SoV)
Lokrix AI computes share of voice as the presence-ratio SoV: p̃_target / (p̃_target + Σp̃_competitors), where p̃ is the Beta-shrunk presence probability per competitor. A position-weighted variant uses PAWC (1/log2(rank+1)) so a #1 citation is not treated the same as a #5 citation.
Sequential Wilson stopping
The adaptive sampling rule that controls how many draws Lokrix AI takes per (prompt × engine × cell). It samples while the Wilson 95% CI width exceeds 2ε (default ε=0.10), with a minimum of 8 and maximum of 60 draws. Decisive cells (presence probability > 0.8 or < 0.2) exit early to save cost. This concentrates draws exactly where uncertainty is high.
Mode separation
The rule that only ORGANIC draws (recommendation-elicitation, intent fan-out) enter the headline composite. DIAGNOSTIC-ELICITED draws (citation-forced, brand-named, comparison) are reported separately. Blending a diagnostic mode — where the brand is named in the prompt — into the organic presence probability is the cardinal statistical-honesty violation.
Provenance
A badge on every data point indicating the measurement method: "live" (a real-time engine run), "cached" (a stored result from a recent run), or "predicted" (the calibrated ML model). Live probabilities carry an "uncalibrated" flag until the Platt/Isotonic calibration map has been fit from historical predicted→realized outcomes. Provenance is always visible.

See your probability for the first time.

Start free — 50 trial credits, no credit card — and run your prompt universe automatically across seven engines — Grok in beta. Add a plan plus credits to unlock live Monte-Carlo sampling, history, regression alerts, and the optimization engine.

Start free · 50 trial credits · no credit card · then a plan + credits (1 credit = 1 live engine response)