AI搜索结果不固定,单次测量会误判领域影响力。
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
- 将引用可见性视为样本估计,而非固定值
- 多轮采样显示引用分布呈幂律,排名极不稳定
- 需报告置信区间,避免误读领域表现差异
AI驱动的问答引擎具有内在随机性:相同查询在不同时间可能产生不同回答和引用来源。当前衡量生成式搜索中领域可见性的方法通常依赖单次运行的引用占比和出现频率点估计,隐含将其视为固定值。本文认为,引用可见性指标应被视为底层响应分布的样本估计量,而非固定数值。我们通过在三个生成式搜索平台(Perplexity Search、OpenAI SearchGPT、Google Gemini)上对三个消费者产品主题进行重复采样,开展实证研究。采样方式包括九天每日采集和每十分钟高频采样。结果显示,引用分布呈幂律形式,且在多次采样间存在显著变异性。基于自助法的置信区间表明,许多看似显著的领域差异实际落在测量噪声范围内。全分布排名稳定性分析进一步显示,不仅头部领域,整个常被引用领域集的排名在不同样本间均不稳定。这些发现表明,单次运行的可见性指标会给人以错误的精确感。我们主张,引用可见性必须附带不确定性估计,并提供实现可解释置信区间的样本量建议。
原文摘要 · Abstract (English)
AI-powered answer engines are inherently non-deterministic: identical queries submitted at different times can produce different responses and cite different sources. Despite this stochastic behavior, current approaches to measuring domain visibility in generative search typically rely on single-run point estimates of citation share and prevalence, implicitly treating them as fixed values. This paper argues that citation visibility metrics should be treated as sample estimators of an underlying response distribution rather than fixed values. We conduct an empirical study of citation variability across three generative search platforms--Perplexity Search, OpenAI SearchGPT, and Google Gemini--using repeated sampling across three consumer product topics. Two sampling regimes are employed: daily collections over nine days and high-frequency sampling at ten-minute intervals. We show that citation distributions follow a power-law form and exhibit substantial variability across repeated samples. Bootstrap confidence intervals reveal that many apparent differences between domains fall within the noise floor of the measurement process. Distribution-wide rank stability analysis further demonstrates that citation rankings are unstable across samples, not only among top-ranked domains but throughout the frequently cited domain set. These findings demonstrate that single-run visibility metrics provide a misleadingly precise picture of domain performance in generative search. We argue that citation visibility must be reported with uncertainty estimates and provide practical guidance for sample sizes required to achieve interpretable confidence intervals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。