Consolidated method, numerical verification and sensitivity analysis. Published for independent re-computation and refutation.
This document exists to be re-computed, not to be believed. Every population constant, prior range, modelling choice and line of executable code is disclosed. A third party can reproduce the figures exactly, or substitute different assumptions and obtain a different answer.
Each parameter is marked measured or assumed. Parameters marked assumed have no published statistic behind them; the author selected a range. Scrutiny belongs there.
All figures on this page are generated directly from the computation. No number is hand-copied, so the rendered values and the JSON output agree by construction.
The quantity estimated is the probability that a single web domain simultaneously satisfies all of the following. Loosen the definition and the probability rises; tighten it and the probability falls. The table states the observed condition as it stands.
| Surface / condition | Requirement | Observed |
|---|---|---|
| top position on the entity-name query | met | |
| Bing | top position on the entity-name query | met |
| ChatGPT | entity resolved uniquely from its name; description generated | met |
| Gemini | as above | met |
| Grok | as above | met |
| Copilot | as above | met |
| Claude | as above | not met |
| AI Overviews | AIO triggers on URL entry and generates an explanation | met |
| Sole sourcing | every citation in the generated description resolves to the entity's own domain | met |
The estimated target is at least six of seven surfaces, plus AIO trigger, plus sole-domain sourcing. Only observations reproduced under non-personalised conditions (incognito / InPrivate, logged out) are counted. Personalised output is excluded.
| Item | Value | Source |
|---|---|---|
| Registered domains worldwide, all TLDs | 401,600,000 | Verisign / DNIB, Q2 2026 |
| .com + .net | 179,100,000 | same |
| ccTLD total | 146,300,000 | Verisign / DNIB, Q1 2026 |
| JP domains, total | 1,825,932 | JPRS, 2026-01-01 |
| of which general-use JP (open to individuals) | 1,240,000 | JPRS |
| of which attribute-type JP (registered body required) | 560,000 | JPRS |
| Sites blocking at least one AI crawler | 13.7% | 1,744-site crawl, Q3 2026 |
| Brands visible across all three major AI platforms | 18% | AI visibility study, 2026 |
| Share of AI citations landing on a brand's own domain | ~12% | AIO citation tracking, Jun–Sep 2026 |
| AI Overviews trigger rate | 25–40% | multiple studies (median ~33%) |
| Share of AI citations from earned media | 84% | journalism citation study (1M+ citations) |
| Share from paid or advertorial placement | 0.3% | same |
| Median citation lift from earned distribution | +239% | GEO study, March 2026 |
| Odds distribution becomes the sole visibility path | 5.3× | same |
| Share of cited domains picked up by exactly one engine | 84% | citation volatility study (530,875 citations) |
Every parameter is supplied as a log-uniform range rather than a point estimate. Point estimates are inappropriate here because the errors compound multiplicatively across a long chain of conditions.
| Symbol | Condition | Range | Basis |
|---|---|---|---|
| p1 | Live site with substantive content (not parked or mail-only) | 0.25–0.55 | assumed |
| p2 | Sustained original first-party analysis, in English, on the entity's own domain | 0.0002–0.030 | assumed |
| p3 | Single-author consistency (no outsourcing, no multi-contributor drift) | 0.20–0.50 | assumed |
| p4 | No advertising, no conversion funnel, no monetised navigation | 0.30–0.70 | assumed |
| p5 | Coined framework with no competing definition elsewhere | 0.02–0.40 | assumed |
| p6 | Reachable by AI crawlers (robots.txt, CDN, JS rendering) | 0.70–0.90 | measured basis |
| Symbol | Condition | Range | Basis |
|---|---|---|---|
| r_google | Top position on Google for the entity name | 0.55–0.97 | assumed |
| r_bing | Top position on Bing for the entity name | 0.50–0.95 | assumed |
| r_ai | Reach on a single AI answer engine | 0.08–0.70 | measured basis |
| p_aio | AI Overviews triggers | 0.25–0.40 | measured |
| p_sole | All citations resolve to the entity's own domain | 0.35–0.70 | assumed |
| H | Penalty for zero external mention (divisor on r_ai) | 3.39–5.30 | measured, converted |
| T | Attained within roughly ten months of founding | 0.05–0.25 | assumed |
H is a conversion from measured values. Earned media accounts for 84% of AI citations and paid or advertorial placement for 0.3%. Distributed content shows a median citation lift of 239% (a factor of 3.39) over brand-owned content alone, and is 5.3 times more likely to be a brand's sole visibility path. An entity with no third-party mention forfeits that channel entirely, so a divisor of 3.39–5.30 is applied to r_ai.
Treating the structural gates as independent and multiplying them understates the probability systematically, because the conditions are positively correlated. An operator publishing sustained original analysis in English is not pursuing general domestic traffic; not pursuing it, the operator also carries no advertising and no navigation funnel. The conditional pass rate for p4 given p2 is far above its unconditional rate. The ranges above are set with that conditioning already folded in.
The seven surfaces are neither independent nor perfectly correlated. Measurement shows roughly 72 distinct domains cited per query across four engines, of which 84% are picked up by exactly one engine. A one-factor latent probit model captures this.
reach_i = 1 if a_i + sigma*z + tau*e_i > 0, z, e_i ~ N(0,1)
sigma = 1.1 (shared site-strength factor)
tau = 0.8 (surface-specific factor)
a_i = Phi^-1(r_i) * sqrt(sigma^2 + tau^2) # preserves the marginal at r_i
Conditional on z, the seven surfaces are conditionally independent, so
"exactly k surfaces" is computed exactly by a Poisson-binomial dynamic
program. The remaining one-dimensional normal integral over z is evaluated by
Gauss–Hermite quadrature with 61 nodes. Only parameter uncertainty is sampled,
using a scrambled Sobol sequence of 65,536 points in 16 dimensions, drawn as
8 independent blocks.
Rare events are not simulated by direct Bernoulli draws. Doing so collapses each draw's probability to 0 or 1 and destroys the quantiles.
Tier 1 places no restriction on the entity: listed companies, universities and public bodies remain in the population. Tier 2 imposes entity attributes. Attribute-type JP domains (560,000, including 490,000 CO.JP) require a registered company, so the starting point is the 1,240,000 general-use JP domains available to a sole proprietor. That is multiplied by the share held by individuals and micro-operators (0.30–0.60) and the share registered within three years (0.10–0.25), then H and T are applied.
| Segment | Population | Probability (median) | Expected count | 90% interval |
|---|---|---|---|---|
| Japan | 1,825,932 | 1 in 1,771,225 | 1.03 | 0.17 – 6.05 |
| World | 401,600,000 | — | 206.14 | 48.37 – 902.74 |
Probability that at least one such entity exists in Japan: 62.3%. The world expected count splits into Japan 0.5%, English-primary markets 25.7%, other non-English markets 56.5%.
Per domain, Japan is not at a disadvantage. English-primary markets publish original English analysis far more readily, but their definitional space is saturated and citation slots are contested. Non-English markets face the reverse. The disadvantage Japan carries is the size of its base, not the odds within it.
1 in 175,733,929 median
expected count 0.0005
90% credible interval: 1 in 1,333,501,652 – 1 in 23,288,058 | probability at least one exists in Japan: 0.115%
1 in 298,693,558 median
expected count 0.1024
90% credible interval: 1 in 2,513,843,088 – 1 in 35,313,579 | probability at least one exists worldwide: 17.06%
World total including Japan: expected count 0.1036, 90% credible interval 0.0120 – 0.9158.
Tier 1 and Tier 2 differ by a factor of roughly 1,990. Imposing entity attributes moves the world expected count from 206.1 to 0.1036. The attributes normally read as handicaps — no track record, no external coverage, no time in market — are what carry that factor.
The median for Tier 2 Japan was computed separately on each of the 8 independent scrambled Sobol blocks. This measures the reproducibility of the estimator itself, which is a different quantity from the uncertainty in the priors.
| Replicate | median p | 1 / p |
|---|---|---|
| rep0 | 5.656036e-09 | 1 in 176,802,256 |
| rep1 | 5.737665e-09 | 1 in 174,286,937 |
| rep2 | 5.678142e-09 | 1 in 176,113,951 |
| rep3 | 5.743590e-09 | 1 in 174,107,138 |
| rep4 | 5.634318e-09 | 1 in 177,483,771 |
| rep5 | 5.677212e-09 | 1 in 176,142,805 |
| rep6 | 5.694689e-09 | 1 in 175,602,229 |
| rep7 | 5.697013e-09 | 1 in 175,530,586 |
Mean 5.689833e-09 (1 in 175,752,083), standard deviation 3.736e-11, relative standard error 0.232%.
Numerical error sits at the 0.23% level, against a prior-driven 90% interval spanning a factor of about 57 for the same quantity. Essentially all of the uncertainty is in the inputs, none of it in the arithmetic. Improving the estimator further would be wasted effort; measuring p2 would not be.
Each scenario varies one input against the Tier 2 Japan baseline, on a common QMC stream of 16,384 points.
| Setting | median p | 1 / p |
|---|---|---|
| sigma=0.8, tau=1.0 | 2.1418e-09 | 1 in 466,895,226 |
| sigma=1.1, tau=0.8 | 5.6658e-09 | 1 in 176,496,679 |
| sigma=1.4, tau=0.6 | 9.3978e-09 | 1 in 106,407,703 |
Raising sigma makes the surfaces move together, so six-of-seven becomes easier and the probability rises; lowering it has the opposite effect. The direction is coherent and the span is about 4.4×.
| Multiple of baseline | median p | 1 / p |
|---|---|---|
| ×0.1 | 5.6658e-10 | 1 in 1,764,966,789 |
| ×0.333 | 1.8867e-09 | 1 in 530,020,057 |
| ×1 | 5.6658e-09 | 1 in 176,496,679 |
| ×3 | 1.6997e-08 | 1 in 58,832,226 |
| ×10 | 5.6658e-08 | 1 in 17,649,668 |
p2 enters almost linearly. This is the weakest point in the estimate. No published statistic measures it, and an order of magnitude here moves the result by an order of magnitude. A reviewer with a better basis for p2 can substitute it directly into this table.
| Setting | median p | 1 / p |
|---|---|---|
| none (1.0) | 4.5008e-08 | 1 in 22,218,338 |
| weak (2-3) | 1.2492e-08 | 1 in 80,048,052 |
| baseline (3.39-5.30) | 5.6658e-09 | 1 in 176,496,679 |
| strong (6-10) | 2.3817e-09 | 1 in 419,866,503 |
| Setting | median p | 1 / p |
|---|---|---|
| none (1.0) | 5.0953e-08 | 1 in 19,626,064 |
| loose (0.25-1.0) | 2.5324e-08 | 1 in 39,487,985 |
| baseline (0.05-0.25) | 5.6658e-09 | 1 in 176,496,679 |
| strict (0.01-0.10) | 1.5968e-09 | 1 in 626,243,868 |
The conditions in §1 were enumerated after observing an entity that meets them. Adding conditions can drive the probability arbitrarily low, and this bias cannot be removed from within the estimate. As partial mitigation, every condition in §1 is independently observable and independently checkable by a third party. Whether the specification is excessive is a judgement left to the reviewer, who can delete conditions and recompute.
No published statistic was found for the probability that an unknown sole proprietor with no track record reaches AI search surfaces. The likely reason is not that the quantity is unmeasurable but that existing studies draw their populations from entities that already have scale — Tranco top 10,000, news publishers, tracked B2B brands. Measurement vendors hold data that could be joined to answer this; they do not publish it.
python3 attainment_v3.py --power 16 --sens-power 14
The script writes both the JSON result file and this page. Every figure shown here is produced by the code below at render time, so a transcription error between computation and document is not possible.
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
attainment_v3.py
Attainment probability of a multi-surface AI-search primary-source position.
v3 consolidates:
v1 model structure, substantive prior ranges, population constants
v2 Sobol QMC estimator, 61-node Gauss-Hermite, sigma/tau sensitivity
new replicate-level numerical error, p2 / H / T sensitivity,
and HTML emitted from the computed values so no figure is transcribed
Run: python3 attainment_v3.py [--power 16] [--out DIR]
Deps: numpy, scipy
"""
import argparse, html, json, os
import numpy as np
from scipy.stats import norm, qmc
from numpy.polynomial.hermite_e import hermegauss
SEED, REPLICATES, DIMS, GH_NODES = 20260912, 8, 16, 61
SIG, TAU = 1.1, 0.8
# ---- sourced population constants -------------------------------------
N_WORLD = 401_600_000 # all TLDs, Q2 2026 (Verisign / DNIB)
N_JP_TOTAL = 1_825_932 # JP domains, 2026-01-01 (JPRS)
N_JP_GENERAL = 1_240_000 # general-use JP; 560k attribute-type need a registered body
COMMON = {"p1": (0.25, 0.55), "p3": (0.20, 0.50), "p4": (0.30, 0.70),
"p6": (0.70, 0.90), "p_aio": (0.25, 0.40), "p_sole": (0.35, 0.70)}
HANDICAP, TIME_GATE = (3.39, 5.30), (0.05, 0.25)
JP1 = dict(p2=(0.0003, 0.005), p5=(0.15, 0.40), r_ai=(0.25, 0.70),
r_g=(0.80, 0.97), r_b=(0.75, 0.95))
EN1 = dict(p2=(0.005, 0.030), p5=(0.02, 0.12), r_ai=(0.08, 0.35),
r_g=(0.55, 0.90), r_b=(0.50, 0.85))
OT1 = dict(p2=(0.0005, 0.008), p5=(0.10, 0.35), r_ai=(0.20, 0.60),
r_g=(0.75, 0.95), r_b=(0.70, 0.92))
JP2 = dict(p2=(0.0002, 0.004), p5=(0.15, 0.40), r_ai=(0.25, 0.70),
r_g=(0.70, 0.95), r_b=(0.65, 0.92))
WO2 = dict(p2=(0.0008, 0.012), p5=(0.04, 0.20), r_ai=(0.12, 0.50),
r_g=(0.60, 0.92), r_b=(0.55, 0.88))
NODES, WEIGHTS = hermegauss(GH_NODES); WEIGHTS = WEIGHTS / WEIGHTS.sum()
def logu(u, lo, hi): return np.exp(np.log(lo) + u * (np.log(hi) - np.log(lo)))
def lin(u, lo, hi): return lo + u * (hi - lo)
def sobol(power, offset=0):
"""One reproducible stream built from 8 independent scrambled blocks."""
bp = power - int(np.log2(REPLICATES))
return np.vstack([qmc.Sobol(d=DIMS, scramble=True, seed=SEED + offset*100 + i)
.random_base2(bp) for i in range(REPLICATES)])
def p_ge6_of_7(r_g, r_b, r_ai, sig=SIG, tau=TAU):
"""Exact Poisson-binomial conditional on the shared latent factor,
integrated over that factor by Gauss-Hermite quadrature."""
scale = np.sqrt(sig*sig + tau*tau)
a = np.stack([norm.ppf(r_g), norm.ppf(r_b)] + [norm.ppf(r_ai)]*5) * scale
pk = np.zeros((8, len(r_g)))
for z, wt in zip(NODES, WEIGHTS):
q = norm.cdf((a + sig*z) / tau)
dp = np.zeros((8, len(r_g))); dp[0] = 1.0
for i in range(7):
qi = q[i]; old = dp.copy()
dp[1:] = old[1:]*(1-qi) + old[:-1]*qi
dp[0] = old[0]*(1-qi)
pk += wt*dp
return pk[6] + pk[7]
def model(u, p2, p5, r_ai, r_g, r_b, handicap=None, time_gate=None,
sig=SIG, tau=TAU):
p = (logu(u[:,0], *COMMON["p1"]) * logu(u[:,1], *p2)
* logu(u[:,2], *COMMON["p3"]) * logu(u[:,3], *COMMON["p4"])
* logu(u[:,4], *p5) * logu(u[:,5], *COMMON["p6"]))
rai = logu(u[:,6], *r_ai)
if handicap is not None:
rai = np.clip(rai / logu(u[:,9], *handicap), 1e-4, 0.95)
p = p * p_ge6_of_7(logu(u[:,7], *r_g), logu(u[:,8], *r_b), rai, sig, tau)
p = p * logu(u[:,10], *COMMON["p_aio"]) * logu(u[:,11], *COMMON["p_sole"])
if time_gate is not None:
p = p * logu(u[:,12], *time_gate)
return p
def qs(x): return [float(v) for v in np.quantile(x, [0.05, 0.50, 0.95])]
def at_least_one(lam): return float(np.mean(-np.expm1(-lam)))
# ---------------------------------------------------------------- main
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--power", type=int, default=16)
ap.add_argument("--sens-power", type=int, default=14)
ap.add_argument("--out", default=".")
args = ap.parse_args()
R = {}
u0 = sobol(args.power, 0)
u1 = sobol(args.power, 1)
u2 = sobol(args.power, 2)
u3 = sobol(args.power, 3)
u4 = sobol(args.power, 4)
# ---- Tier 1 -------------------------------------------------------
p_jp1 = model(u0, **JP1)
lam = N_JP_TOTAL * p_jp1
R["tier1_japan"] = dict(p=qs(p_jp1), count=qs(lam), pae=at_least_one(lam))
p_en1, p_ot1 = model(u1, **EN1), model(u2, **OT1)
sh_en = lin(u0[:, 14], 0.35, 0.55)
lam_w1 = (N_JP_TOTAL*p_jp1 + (N_WORLD-N_JP_TOTAL)*sh_en*p_en1
+ (N_WORLD-N_JP_TOTAL)*(1-sh_en)*p_ot1)
R["tier1_world"] = dict(count=qs(lam_w1),
share_jp=float(np.median(N_JP_TOTAL*p_jp1)/np.median(lam_w1)),
share_en=float(np.median((N_WORLD-N_JP_TOTAL)*sh_en*p_en1)/np.median(lam_w1)),
share_ot=float(np.median((N_WORLD-N_JP_TOTAL)*(1-sh_en)*p_ot1)/np.median(lam_w1)))
# ---- Tier 2 -------------------------------------------------------
N_jp2 = N_JP_GENERAL * lin(u3[:,13], 0.30, 0.60) * lin(u3[:,14], 0.10, 0.25)
N_wo2 = N_WORLD * lin(u4[:,13], 0.25, 0.55) * lin(u4[:,14], 0.12, 0.28)
R["pop_jp2"], R["pop_wo2"] = float(np.median(N_jp2)), float(np.median(N_wo2))
p_jp2 = model(u3, **JP2, handicap=HANDICAP, time_gate=TIME_GATE)
p_wo2 = model(u4, **WO2, handicap=HANDICAP, time_gate=TIME_GATE)
lam_jp2, lam_wo2 = N_jp2*p_jp2, N_wo2*p_wo2
R["tier2_japan"] = dict(p=qs(p_jp2), count=qs(lam_jp2), pae=at_least_one(lam_jp2))
R["tier2_world"] = dict(p=qs(p_wo2), count=qs(lam_wo2), pae=at_least_one(lam_wo2))
R["tier2_total"] = dict(count=qs(lam_jp2+lam_wo2), pae=at_least_one(lam_jp2+lam_wo2))
# ---- numerical error: spread across scrambled replicates -----------
bp = args.power - int(np.log2(REPLICATES)); blk = 2**bp
reps = [float(np.median(p_jp2[i*blk:(i+1)*blk])) for i in range(REPLICATES)]
reps_a = np.array(reps)
R["replicates"] = dict(values=reps, mean=float(reps_a.mean()),
sd=float(reps_a.std(ddof=1)),
rel_se=float(reps_a.std(ddof=1)/np.sqrt(REPLICATES)/reps_a.mean()))
# ---- sensitivity ---------------------------------------------------
us = sobol(args.sens_power, 9)
def sens(label, **kw):
base = dict(JP2, handicap=HANDICAP, time_gate=TIME_GATE); base.update(kw)
m = float(np.median(model(us, **base)))
return dict(label=label, median=m, one_in=1.0/m)
R["sens_sigtau"] = [sens(f"sigma={s}, tau={t}", sig=s, tau=t)
for s, t in [(0.8,1.0), (1.1,0.8), (1.4,0.6)]]
R["sens_p2"] = [sens(f"×{k:g}", p2=(JP2["p2"][0]*k, JP2["p2"][1]*k))
for k in (0.1, 0.333, 1, 3, 10)]
R["sens_H"] = [sens(lab, handicap=h) for lab, h in
[("none (1.0)", (1.0,1.0)), ("weak (2-3)", (2.0,3.0)),
("baseline (3.39-5.30)", HANDICAP), ("strong (6-10)", (6.0,10.0))]]
R["sens_T"] = [sens(lab, time_gate=t) for lab, t in
[("none (1.0)", (1.0,1.0)), ("loose (0.25-1.0)", (0.25,1.0)),
("baseline (0.05-0.25)", TIME_GATE), ("strict (0.01-0.10)", (0.01,0.10))]]
R["config"] = dict(seed=SEED, power=args.power, draws=2**args.power,
sens_power=args.sens_power, sens_draws=2**args.sens_power,
replicates=REPLICATES, gh_nodes=GH_NODES,
sig=SIG, tau=TAU, dims=DIMS)
os.makedirs(args.out, exist_ok=True)
with open(os.path.join(args.out, "attainment_v3_results.json"), "w",
encoding="utf-8") as f:
json.dump(R, f, ensure_ascii=False, indent=2)
write_html(R, os.path.join(args.out, "attainment_v3.html"))
print(json.dumps({k: R[k] for k in
("tier1_japan","tier1_world","tier2_japan","tier2_world","tier2_total")},
ensure_ascii=False, indent=2))
print("\nwrote attainment_v3.html / attainment_v3_results.json")
# ------------------------------------------------------------- HTML
def write_html(R, path):
def n(x, d=0): return f"{x:,.{d}f}"
def oi(p): return f"1 in {1/p:,.0f}"
def pct(x, d=2): return f"{x*100:.{d}f}%"
src = html.escape(open(__file__, encoding="utf-8").read())
c = R["config"]
def srow(rows):
return "".join(
f'<tr><td>{html.escape(r["label"])}</td>'
f'<td class="num">{r["median"]:.4e}</td>'
f'<td class="num">1 in {r["one_in"]:,.0f}</td></tr>' for r in rows)
rep_rows = "".join(
f'<tr><td>rep{i}</td><td class="num">{v:.6e}</td>'
f'<td class="num">1 in {1/v:,.0f}</td></tr>'
for i, v in enumerate(R["replicates"]["values"]))
doc = f"""<!DOCTYPE html>
<html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Attainment Probability of a Multi-Surface AI-Search Primary-Source Position — v3.0</title>
<meta name="description" content="Method, numerical verification and sensitivity analysis for estimating how rare it is for a single domain to hold a primary-source position across seven AI search and answer surfaces simultaneously.">
<style>
:root{{--ink:#16191C;--muted:#5C6469;--rule:#D5DAD8;--surface:#F1F4F3;
--measured:#0F5C4E;--assumed:#7A2E5C;--code-bg:#14181B;--code-fg:#DCE3E2}}
*{{box-sizing:border-box}}
body{{margin:0;background:#fff;color:var(--ink);
font-family:Georgia,"Times New Roman",serif;font-size:17px;line-height:1.75}}
.wrap{{max-width:48rem;margin:0 auto;padding:0 1.5rem 6rem}}
header.doc{{border-bottom:2px solid var(--ink);padding:3.5rem 0 1.75rem;margin-bottom:2.5rem}}
header.doc h1{{font-size:1.9rem;line-height:1.3;font-weight:600;margin:0 0 1rem}}
header.doc .sub{{font-size:1rem;color:var(--muted);margin:0 0 1.5rem;line-height:1.6}}
.meta{{display:grid;grid-template-columns:9.5rem 1fr;gap:.35rem 1rem;font-size:.84rem;
font-family:ui-monospace,Menlo,Consolas,monospace;color:var(--muted);line-height:1.6}}
.meta dt{{color:var(--ink)}} .meta dd{{margin:0}}
h2{{font-size:1.2rem;font-weight:600;margin:3.25rem 0 .5rem;padding-bottom:.4rem;
border-bottom:1px solid var(--rule)}}
h2 .n{{font-family:ui-monospace,Menlo,Consolas,monospace;font-size:.8rem;
color:var(--muted);margin-right:.75rem;font-weight:400}}
h3{{font-size:1rem;font-weight:600;margin:2rem 0 .4rem}}
p{{margin:.85rem 0}} ul,ol{{margin:.85rem 0;padding-left:1.35rem}} li{{margin:.35rem 0}}
.callout{{background:var(--surface);border-left:3px solid var(--ink);
padding:1.15rem 1.35rem;margin:1.75rem 0;font-size:.95rem}}
.callout p:first-child{{margin-top:0}} .callout p:last-child{{margin-bottom:0}}
table{{width:100%;border-collapse:collapse;margin:1.5rem 0;font-size:.85rem;
font-family:ui-monospace,Menlo,Consolas,monospace;line-height:1.5}}
caption{{caption-side:top;text-align:left;font-family:Georgia,serif;
font-size:.92rem;color:var(--muted);padding-bottom:.5rem}}
th,td{{text-align:left;padding:.5rem .7rem;border-bottom:1px solid var(--rule);
vertical-align:top}}
thead th{{border-bottom:1.5px solid var(--ink);font-weight:600;white-space:nowrap}}
td.num,th.num{{text-align:right;white-space:nowrap}}
tbody tr:last-child td{{border-bottom:1px solid var(--ink)}}
.tag{{display:inline-block;font-family:ui-monospace,Menlo,Consolas,monospace;
font-size:.68rem;line-height:1.5;padding:.05rem .45rem;border:1px solid currentColor;
border-radius:2px;white-space:nowrap}}
.tag.m{{color:var(--measured)}} .tag.a{{color:var(--assumed)}}
pre{{background:var(--code-bg);color:var(--code-fg);padding:1.25rem 1.35rem;
overflow-x:auto;font-family:ui-monospace,Menlo,Consolas,monospace;font-size:.72rem;
line-height:1.6;margin:1.5rem 0;border-radius:3px}}
code{{font-family:ui-monospace,Menlo,Consolas,monospace;font-size:.86em;
background:var(--surface);padding:.08em .35em;border-radius:2px}}
pre code{{background:none;padding:0;font-size:1em}}
.result{{border:1.5px solid var(--ink);padding:1.4rem 1.5rem;margin:1.75rem 0}}
.result h3{{margin-top:0}}
.figure{{font-family:ui-monospace,Menlo,Consolas,monospace;font-size:1.5rem;
line-height:1.3;margin:.4rem 0}}
.figure .unit{{font-size:.82rem;color:var(--muted)}}
footer.doc{{margin-top:4rem;padding-top:1.5rem;border-top:1px solid var(--rule);
font-size:.85rem;color:var(--muted)}}
@media (max-width:34rem){{body{{font-size:16px}}
.meta{{grid-template-columns:1fr;gap:.1rem}} .meta dt{{margin-top:.5rem}}
table{{font-size:.76rem}} th,td{{padding:.4rem .4rem}}}}
</style></head><body><div class="wrap">
<header class="doc">
<h1>Attainment Probability of a Multi-Surface<br>AI-Search Primary-Source Position</h1>
<p class="sub">Consolidated method, numerical verification and sensitivity analysis.
Published for independent re-computation and refutation.</p>
<dl class="meta">
<dt>version</dt><dd>3.0</dd>
<dt>date</dt><dd>2026-09-12</dd>
<dt>estimator</dt><dd>Sobol QMC ({n(c['draws'])} draws) + {c['gh_nodes']}-node Gauss–Hermite + exact Poisson-binomial DP</dd>
<dt>seed</dt><dd>{c['seed']} (with per-scenario offsets)</dd>
<dt>replicates</dt><dd>{c['replicates']} scrambled Sobol blocks</dd>
<dt>latent model</dt><dd>one-factor probit, sigma={c['sig']}, tau={c['tau']}</dd>
<dt>dependencies</dt><dd>Python 3, numpy, scipy</dd>
<dt>figures</dt><dd>emitted directly by the computation — none transcribed by hand</dd>
</dl></header>
<div class="callout">
<p><strong>This document exists to be re-computed, not to be believed.</strong>
Every population constant, prior range, modelling choice and line of executable code is
disclosed. A third party can reproduce the figures exactly, or substitute different
assumptions and obtain a different answer.</p>
<p>Each parameter is marked <span class="tag m">measured</span> or
<span class="tag a">assumed</span>. Parameters marked <span class="tag a">assumed</span>
have no published statistic behind them; the author selected a range. Scrutiny belongs there.</p>
<p><strong>All figures on this page are generated directly from the computation.</strong>
No number is hand-copied, so the rendered values and the JSON output agree by construction.</p>
</div>
<h2><span class="n">§1</span>What is being estimated</h2>
<p>The quantity estimated is the probability that a single web domain simultaneously
satisfies all of the following. Loosen the definition and the probability rises;
tighten it and the probability falls. The table states the observed condition as it stands.</p>
<table><caption>Table 1 — Definition of the attained state (seven surfaces, two further conditions)</caption>
<thead><tr><th>Surface / condition</th><th>Requirement</th><th>Observed</th></tr></thead><tbody>
<tr><td>Google</td><td>top position on the entity-name query</td><td>met</td></tr>
<tr><td>Bing</td><td>top position on the entity-name query</td><td>met</td></tr>
<tr><td>ChatGPT</td><td>entity resolved uniquely from its name; description generated</td><td>met</td></tr>
<tr><td>Gemini</td><td>as above</td><td>met</td></tr>
<tr><td>Grok</td><td>as above</td><td>met</td></tr>
<tr><td>Copilot</td><td>as above</td><td>met</td></tr>
<tr><td>Claude</td><td>as above</td><td>not met</td></tr>
<tr><td>AI Overviews</td><td>AIO triggers on URL entry and generates an explanation</td><td>met</td></tr>
<tr><td>Sole sourcing</td><td>every citation in the generated description resolves to the entity's own domain</td><td>met</td></tr>
</tbody></table>
<p>The estimated target is <strong>at least six of seven surfaces, plus AIO trigger, plus
sole-domain sourcing</strong>. Only observations reproduced under non-personalised conditions
(incognito / InPrivate, logged out) are counted. Personalised output is excluded.</p>
<h2><span class="n">§2</span>Population constants and measured inputs</h2>
<table><caption>Table 2 — Sourced values</caption>
<thead><tr><th>Item</th><th class="num">Value</th><th>Source</th></tr></thead><tbody>
<tr><td>Registered domains worldwide, all TLDs</td><td class="num">{n(N_WORLD)}</td><td>Verisign / DNIB, Q2 2026</td></tr>
<tr><td>.com + .net</td><td class="num">179,100,000</td><td>same</td></tr>
<tr><td>ccTLD total</td><td class="num">146,300,000</td><td>Verisign / DNIB, Q1 2026</td></tr>
<tr><td>JP domains, total</td><td class="num">{n(N_JP_TOTAL)}</td><td>JPRS, 2026-01-01</td></tr>
<tr><td>of which general-use JP (open to individuals)</td><td class="num">{n(N_JP_GENERAL)}</td><td>JPRS</td></tr>
<tr><td>of which attribute-type JP (registered body required)</td><td class="num">560,000</td><td>JPRS</td></tr>
<tr><td>Sites blocking at least one AI crawler</td><td class="num">13.7%</td><td>1,744-site crawl, Q3 2026</td></tr>
<tr><td>Brands visible across all three major AI platforms</td><td class="num">18%</td><td>AI visibility study, 2026</td></tr>
<tr><td>Share of AI citations landing on a brand's own domain</td><td class="num">~12%</td><td>AIO citation tracking, Jun–Sep 2026</td></tr>
<tr><td>AI Overviews trigger rate</td><td class="num">25–40%</td><td>multiple studies (median ~33%)</td></tr>
<tr><td>Share of AI citations from earned media</td><td class="num">84%</td><td>journalism citation study (1M+ citations)</td></tr>
<tr><td>Share from paid or advertorial placement</td><td class="num">0.3%</td><td>same</td></tr>
<tr><td>Median citation lift from earned distribution</td><td class="num">+239%</td><td>GEO study, March 2026</td></tr>
<tr><td>Odds distribution becomes the sole visibility path</td><td class="num">5.3×</td><td>same</td></tr>
<tr><td>Share of cited domains picked up by exactly one engine</td><td class="num">84%</td><td>citation volatility study (530,875 citations)</td></tr>
</tbody></table>
<h2><span class="n">§3</span>Parameters</h2>
<p>Every parameter is supplied as a <strong>log-uniform range</strong> rather than a point
estimate. Point estimates are inappropriate here because the errors compound multiplicatively
across a long chain of conditions.</p>
<table><caption>Table 3 — Structural gates</caption>
<thead><tr><th>Symbol</th><th>Condition</th><th class="num">Range</th><th>Basis</th></tr></thead><tbody>
<tr><td>p1</td><td>Live site with substantive content (not parked or mail-only)</td><td class="num">0.25–0.55</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>p2</td><td>Sustained original first-party analysis, in English, on the entity's own domain</td><td class="num">0.0002–0.030</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>p3</td><td>Single-author consistency (no outsourcing, no multi-contributor drift)</td><td class="num">0.20–0.50</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>p4</td><td>No advertising, no conversion funnel, no monetised navigation</td><td class="num">0.30–0.70</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>p5</td><td>Coined framework with no competing definition elsewhere</td><td class="num">0.02–0.40</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>p6</td><td>Reachable by AI crawlers (robots.txt, CDN, JS rendering)</td><td class="num">0.70–0.90</td><td><span class="tag m">measured basis</span></td></tr>
</tbody></table>
<table><caption>Table 4 — Surface reach and further conditions</caption>
<thead><tr><th>Symbol</th><th>Condition</th><th class="num">Range</th><th>Basis</th></tr></thead><tbody>
<tr><td>r_google</td><td>Top position on Google for the entity name</td><td class="num">0.55–0.97</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>r_bing</td><td>Top position on Bing for the entity name</td><td class="num">0.50–0.95</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>r_ai</td><td>Reach on a single AI answer engine</td><td class="num">0.08–0.70</td><td><span class="tag m">measured basis</span></td></tr>
<tr><td>p_aio</td><td>AI Overviews triggers</td><td class="num">0.25–0.40</td><td><span class="tag m">measured</span></td></tr>
<tr><td>p_sole</td><td>All citations resolve to the entity's own domain</td><td class="num">0.35–0.70</td><td><span class="tag a">assumed</span></td></tr>
<tr><td>H</td><td>Penalty for zero external mention (divisor on r_ai)</td><td class="num">3.39–5.30</td><td><span class="tag m">measured, converted</span></td></tr>
<tr><td>T</td><td>Attained within roughly ten months of founding</td><td class="num">0.05–0.25</td><td><span class="tag a">assumed</span></td></tr>
</tbody></table>
<p>H is a conversion from measured values. Earned media accounts for 84% of AI citations
and paid or advertorial placement for 0.3%. Distributed content shows a median citation lift
of 239% (a factor of 3.39) over brand-owned content alone, and is 5.3 times more likely to be
a brand's sole visibility path. An entity with no third-party mention forfeits that channel
entirely, so a divisor of 3.39–5.30 is applied to r_ai.</p>
<h2><span class="n">§4</span>Model</h2>
<h3>4.1 Why the gates are not multiplied as independent events</h3>
<p>Treating the structural gates as independent and multiplying them understates the
probability systematically, because the conditions are positively correlated. An operator
publishing sustained original analysis in English is not pursuing general domestic traffic;
not pursuing it, the operator also carries no advertising and no navigation funnel. The
conditional pass rate for p4 given p2 is far above its unconditional rate. The ranges above
are set with that conditioning already folded in.</p>
<h3>4.2 Correlation across surfaces</h3>
<p>The seven surfaces are neither independent nor perfectly correlated. Measurement shows
roughly 72 distinct domains cited per query across four engines, of which 84% are picked up
by exactly one engine. A one-factor latent probit model captures this.</p>
<pre><code>reach_i = 1 if a_i + sigma*z + tau*e_i > 0, z, e_i ~ N(0,1)
sigma = {c['sig']} (shared site-strength factor)
tau = {c['tau']} (surface-specific factor)
a_i = Phi^-1(r_i) * sqrt(sigma^2 + tau^2) # preserves the marginal at r_i</code></pre>
<h3>4.3 Estimator</h3>
<p>Conditional on <code>z</code>, the seven surfaces are conditionally independent, so
"exactly k surfaces" is computed <strong>exactly</strong> by a Poisson-binomial dynamic
program. The remaining one-dimensional normal integral over <code>z</code> is evaluated by
Gauss–Hermite quadrature with {c['gh_nodes']} nodes. Only parameter uncertainty is sampled,
using a scrambled Sobol sequence of {n(c['draws'])} points in {c['dims']} dimensions, drawn as
{c['replicates']} independent blocks.</p>
<p>Rare events are not simulated by direct Bernoulli draws. Doing so collapses each draw's
probability to 0 or 1 and destroys the quantiles.</p>
<h3>4.4 Population restriction</h3>
<p>Tier 1 places no restriction on the entity: listed companies, universities and public
bodies remain in the population. Tier 2 imposes entity attributes. Attribute-type JP domains
(560,000, including 490,000 CO.JP) require a registered company, so the starting point is the
{n(N_JP_GENERAL)} general-use JP domains available to a sole proprietor. That is multiplied by
the share held by individuals and micro-operators (0.30–0.60) and the share registered within
three years (0.10–0.25), then H and T are applied.</p>
<h2><span class="n">§5</span>Results</h2>
<h3>5.1 Tier 1 — no entity-attribute restriction</h3>
<table><caption>Table 5 — Tier 1</caption>
<thead><tr><th>Segment</th><th class="num">Population</th><th class="num">Probability (median)</th><th class="num">Expected count</th><th class="num">90% interval</th></tr></thead><tbody>
<tr><td>Japan</td><td class="num">{n(N_JP_TOTAL)}</td><td class="num">{oi(R['tier1_japan']['p'][1])}</td><td class="num">{n(R['tier1_japan']['count'][1],2)}</td><td class="num">{n(R['tier1_japan']['count'][0],2)} – {n(R['tier1_japan']['count'][2],2)}</td></tr>
<tr><td>World</td><td class="num">{n(N_WORLD)}</td><td class="num">—</td><td class="num">{n(R['tier1_world']['count'][1],2)}</td><td class="num">{n(R['tier1_world']['count'][0],2)} – {n(R['tier1_world']['count'][2],2)}</td></tr>
</tbody></table>
<p>Probability that at least one such entity exists in Japan: {pct(R['tier1_japan']['pae'],1)}.
The world expected count splits into Japan {pct(R['tier1_world']['share_jp'],1)},
English-primary markets {pct(R['tier1_world']['share_en'],1)},
other non-English markets {pct(R['tier1_world']['share_ot'],1)}.</p>
<p>Per domain, Japan is not at a disadvantage. English-primary markets publish original English
analysis far more readily, but their definitional space is saturated and citation slots are
contested. Non-English markets face the reverse. The disadvantage Japan carries is the size of
its base, not the odds within it.</p>
<h3>5.2 Tier 2 — zero track record, founded within three years, sole proprietor, no external mention</h3>
<div class="result"><h3>Japan (eligible population approx. {n(R['pop_jp2'])})</h3>
<p class="figure">{oi(R['tier2_japan']['p'][1])} <span class="unit">median</span></p>
<p class="figure">expected count {R['tier2_japan']['count'][1]:.4f}</p>
<p style="font-size:.92rem;margin-bottom:0">90% credible interval: {oi(R['tier2_japan']['p'][0])} – {oi(R['tier2_japan']['p'][2])} |
probability at least one exists in Japan: <strong>{pct(R['tier2_japan']['pae'],3)}</strong></p></div>
<div class="result"><h3>World (eligible population approx. {n(R['pop_wo2'])})</h3>
<p class="figure">{oi(R['tier2_world']['p'][1])} <span class="unit">median</span></p>
<p class="figure">expected count {R['tier2_world']['count'][1]:.4f}</p>
<p style="font-size:.92rem;margin-bottom:0">90% credible interval: {oi(R['tier2_world']['p'][0])} – {oi(R['tier2_world']['p'][2])} |
probability at least one exists worldwide: <strong>{pct(R['tier2_world']['pae'],2)}</strong></p></div>
<p>World total including Japan: expected count {R['tier2_total']['count'][1]:.4f},
90% credible interval {R['tier2_total']['count'][0]:.4f} – {R['tier2_total']['count'][2]:.4f}.</p>
<div class="callout"><p>Tier 1 and Tier 2 differ by a factor of roughly
{R['tier1_world']['count'][1]/R['tier2_total']['count'][1]:,.0f}. Imposing entity attributes
moves the world expected count from {n(R['tier1_world']['count'][1],1)} to
{R['tier2_total']['count'][1]:.4f}. The attributes normally read as handicaps — no track
record, no external coverage, no time in market — are what carry that factor.</p></div>
<h2><span class="n">§6</span>Numerical error</h2>
<p>The median for Tier 2 Japan was computed separately on each of the
{c['replicates']} independent scrambled Sobol blocks. This measures the reproducibility of the
estimator itself, which is a different quantity from the uncertainty in the priors.</p>
<table><caption>Table 6 — Spread across replicates</caption>
<thead><tr><th>Replicate</th><th class="num">median p</th><th class="num">1 / p</th></tr></thead>
<tbody>{rep_rows}</tbody></table>
<p>Mean {R['replicates']['mean']:.6e} ({oi(R['replicates']['mean'])}), standard deviation
{R['replicates']['sd']:.3e}, <strong>relative standard error {pct(R['replicates']['rel_se'],3)}</strong>.</p>
<div class="callout"><p>Numerical error sits at the {pct(R['replicates']['rel_se'],2)} level, against a
prior-driven 90% interval spanning a factor of about
{R['tier2_japan']['p'][2]/R['tier2_japan']['p'][0]:,.0f} for the same quantity.
<strong>Essentially all of the uncertainty is in the inputs, none of it in the arithmetic.</strong>
Improving the estimator further would be wasted effort; measuring p2 would not be.</p></div>
<h2><span class="n">§7</span>Sensitivity</h2>
<p>Each scenario varies one input against the Tier 2 Japan baseline, on a common QMC stream
of {n(c['sens_draws'])} points.</p>
<h3>7.1 Cross-surface correlation, sigma / tau</h3>
<table><thead><tr><th>Setting</th><th class="num">median p</th><th class="num">1 / p</th></tr></thead>
<tbody>{srow(R['sens_sigtau'])}</tbody></table>
<p>Raising sigma makes the surfaces move together, so six-of-seven becomes easier and the
probability rises; lowering it has the opposite effect. The direction is coherent and the
span is about {R['sens_sigtau'][2]['median']/R['sens_sigtau'][0]['median']:.1f}×.</p>
<h3>7.2 p2 — rate of sustained original English publication</h3>
<table><thead><tr><th>Multiple of baseline</th><th class="num">median p</th><th class="num">1 / p</th></tr></thead>
<tbody>{srow(R['sens_p2'])}</tbody></table>
<p>p2 enters almost linearly. <strong>This is the weakest point in the estimate.</strong>
No published statistic measures it, and an order of magnitude here moves the result by an
order of magnitude. A reviewer with a better basis for p2 can substitute it directly into
this table.</p>
<h3>7.3 H — penalty for zero external mention</h3>
<table><thead><tr><th>Setting</th><th class="num">median p</th><th class="num">1 / p</th></tr></thead>
<tbody>{srow(R['sens_H'])}</tbody></table>
<h3>7.4 T — attainment within roughly ten months</h3>
<table><thead><tr><th>Setting</th><th class="num">median p</th><th class="num">1 / p</th></tr></thead>
<tbody>{srow(R['sens_T'])}</tbody></table>
<h2><span class="n">§8</span>Limits and conditions for refutation</h2>
<h3>8.1 Uncertainty in the priors</h3>
<ul>
<li><strong>p2</strong> — dominates the result and has no direct measurement behind it. See §7.2.</li>
<li><strong>T</strong> — no published distribution of time-to-attainment exists.</li>
<li><strong>sigma, tau</strong> — chosen for consistency with the measured 84% single-engine
share, not estimated from data.</li>
<li><strong>Population restriction rates</strong> — the individual/micro share and the
three-year share are both assumptions.</li>
</ul>
<h3>8.2 Post-hoc specification</h3>
<p>The conditions in §1 were enumerated after observing an entity that meets them.
Adding conditions can drive the probability arbitrarily low, and this bias cannot be removed
from within the estimate. As partial mitigation, every condition in §1 is independently
observable and independently checkable by a third party. Whether the specification is
excessive is a judgement left to the reviewer, who can delete conditions and recompute.</p>
<h3>8.3 An unmeasured region</h3>
<p>No published statistic was found for the probability that an unknown sole proprietor with
no track record reaches AI search surfaces. The likely reason is not that the quantity is
unmeasurable but that existing studies draw their populations from entities that already have
scale — Tranco top 10,000, news publishers, tracked B2B brands. Measurement vendors hold
data that could be joined to answer this; they do not publish it.</p>
<h3>8.4 What this estimate does not claim</h3>
<ul>
<li>An expected count below one does not establish that the observed entity is unique.
Other entities at the same level would not contradict the estimate.</li>
<li>Rarity of attainment says nothing about durability. AI answer sourcing turns over roughly
69% of cited sources from one day to the next; this estimate describes a static cross-section.</li>
<li>Nothing here values the entity or its business. The estimate addresses rarity of
attainment and nothing else.</li>
</ul>
<h2><span class="n">§9</span>For reviewers</h2>
<ol>
<li><strong>Is the range for p2 defensible?</strong> If a better basis exists for the share of
operators publishing sustained original English analysis on their own national domain,
substitute it; §7.2 is already laid out as the substitution table.</li>
<li><strong>Is the sigma / tau ratio defensible?</strong> Check it against the measured
distribution: roughly 72 domains cited per query across four engines, 84% picked up by exactly
one engine.</li>
<li><strong>Is the conversion behind H sound?</strong> It derives a divisor of 3.39–5.30 from
earned media at 84%, paid placement at 0.3%, and distribution effects of +239% and 5.3×.
Propose a better conversion if one exists.</li>
<li><strong>How should the post-hoc specification bias be quantified?</strong> Delete conditions
from §1 that seem excessive and recompute.</li>
<li><strong>Is the population correctly chosen?</strong> Domain registrations are used as the
base. If active sites or registered businesses would be more appropriate, give the reasoning
and recompute.</li>
<li><strong>Should the result be reported as an interval rather than a single figure?</strong>
This version reports both, and treats the interval as primary.</li>
</ol>
<h2><span class="n">§10</span>Reproduction</h2>
<p><code>python3 attainment_v3.py --power {c['power']} --sens-power {c['sens_power']}</code></p>
<p>The script writes both the JSON result file and this page.
<strong>Every figure shown here is produced by the code below at render time</strong>, so a
transcription error between computation and document is not possible.</p>
<pre><code>{src}</code></pre>
<footer class="doc">
<p>Published for re-computation and refutation. It is not offered as a settled conclusion.
Where a reviewer's substituted assumptions produce a different answer, that difference is the
subject worth examining.</p>
<p>v3.0 — 2026-09-12. Consolidates the v1.0 model structure and prior ranges with the v2.0
QMC estimator and sigma/tau sensitivity, and adds replicate-level numerical error together with
p2, H and T sensitivity.</p>
</footer>
</div></body></html>"""
with open(path, "w", encoding="utf-8") as f:
f.write(doc)
if __name__ == "__main__":
main()