提出新方法判断AI可见性测量何时足够稳定可靠。
From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement

- 基于排名稳定性和结构充分性双标准,动态判定数据收集是否充足。
- 实测30个平台-主题组合,发现固定预算无法通用,需依分布结构判断。
- 无需预设参数,自动适应不同平台与话题,适合做对比分析的团队使用。
AI可见性测量本质上是相对的:从业者希望了解生成式搜索引擎最常引用哪些领域,以及观察到的差异是否足以支撑决策。然而行业缺乏一种系统方法来确定是否已收集足够数据。不同研究和平台的数据采集预算差异巨大,结论常基于稳定性与精确性未知的排名。本文提出一种基于两个互补标准的顺序收敛框架:排名稳定性评估排名相关性轨迹是否达到结构性平稳;结构充分性评估已确立领域(其置信区间不包含零)之间的引文份额差异是否超过估计不确定性。这两个标准均源于观测引文分布的规律,包括其排名结构、不确定性特征及可观测与已确立领域的边界。该框架仅需少量结构性常数,无需外部指定查询次数、相关性目标或置信区间宽度;停止条件由实际测量不确定性驱动,且在多种充分性阈值下保持稳健。在涵盖Gemini、SearchGPT和Perplexity的30个平台-主题组合上应用,框架能自适应不同平台与主题的引文分布。结果表明,无法为所有场景设定统一采集预算,收敛可通过对观测分布结构的分析实现。该框架为判断AI可见性测量是否具备支持比较分析的能力提供了实用依据。
原文摘要 · Abstract (English)
AI visibility measurement is comparative: practitioners want to know which domains generative search engines cite most often and whether observed differences are large enough to support decisions. Yet the industry lacks a principled way to determine whether enough data has been collected. Collection budgets vary widely across studies and platforms, and conclusions are often drawn from rankings whose stability and precision are unknown. We introduce a sequential convergence framework based on two complementary criteria: rank stability evaluates whether the rank-correlation trajectory has reached a structural plateau, while structural sufficiency evaluates whether the spread of citation shares among established domains -- those whose confidence intervals exclude zero -- exceeds the uncertainty of those estimates. Together, these criteria distinguish rankings that have merely stabilized from those sufficiently resolved to support inference. Both are derived from regularities in the observed citation distribution, including its rank structure, uncertainty profile, and the boundary between observed and established domains. The framework retains a small number of structural constants but requires no externally specified query count, correlation target, or confidence-interval width target; stopping is driven by observed measurement uncertainty and remains robust across a range of sufficiency thresholds. Applied across 30 platform-topic combinations spanning Gemini, SearchGPT, and Perplexity, the framework adapts to platform- and topic-specific citation distributions. Results show that no fixed collection budget can be justified across contexts and that convergence can instead be evaluated from the structure of the observed distribution. The framework provides a practical basis for determining when AI visibility measurements are ready to support comparative analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。