同一购买意图的自然改写会导致推荐结果差异巨大,现有追踪方法不可靠。
Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline
- 通过大量改写测试发现,不同问法导致品牌推荐重合度仅14%-29%
- 相同问题重跑的推荐重合度达50%-61%,远高于改写结果
- 当前基于固定问法的可见性追踪方法存在结构性缺陷
买家对问题的微小措辞变化——如'最佳CRM'、'顶级CRM'或'最适合SaaS初创公司的最佳CRM'——会导致AI助手推荐的品牌显著不同。在OpenAI和Anthropic模型上进行约6000次改写实验和同等数量的同提示重跑对照测试显示,相同购买意图的自然改写之间推荐集相似度(Jaccard)仅为0.288(美容类改写,95%置信区间[0.215, 0.361])和0.135(增加约束类改写,[0.098, 0.175]),均远低于相同提示重跑的基准值0.50-0.61。提示字符串而非潜在购买意图成为决定品牌出现的主要因素。增加推理努力无法缩小差距(上下浮动不超过±0.05)。这直接挑战了日益流行的AEO/GEO实践。通过固定提示集统计品牌提及来追踪品牌AI可见性,其主要方差来源是追踪器随机采用的改写方式,而非模型对品牌的实际响应:同一意图的两种自然改写推荐集重合度仅为14%-29%,而同提示重跑为50%-61%。理论上增加每意图的改写样本量可降低偏差,学术界已有高效多提示评估方法,但真实买家提问空间远超这些方法验证的基准规模,也远超商业追踪系统对每个品牌-意图组合发出的提示量。因此,逐提示提及追踪作为测量单位在结构上不稳定;真正改进可能需采用新单位而非更大提示集。
原文摘要 · Abstract (English)
Small changes to how a buyer phrases a question -- "best CRM" vs "top CRM" vs "best CRM for a SaaS startup" -- produce substantially different brand recommendations from AI assistants. Across ~6,000 paraphrase runs and ~6,000 same-prompt rerun controls on OpenAI and Anthropic models, the recommendation-set similarity (Jaccard) between two paraphrases of the same underlying buying intent is 0.288 for cosmetic rewordings (clustered 95% CI [0.215, 0.361]) and 0.135 for constraint-adding rewordings ([0.098, 0.175], pooling region/language and specificity-ladder axes) -- both far below the 0.50-0.61 same-prompt rerun baseline. The prompt string, not the underlying buyer intent, is the dominant input to which brands surface. Increasing reasoning effort does not narrow the gap (bounded by +/-0.05). This is a direct challenge to an increasingly popular AEO/GEO practice. Tracking a brand's "AI visibility" by counting brand mentions over a fixed set of prompts produces a metric whose dominant source of variance is which paraphrase the tracker happens to issue, not the model's behavior toward the brand: the same buyer intent in two natural paraphrases produces recommendation sets that overlap 14-29% in Jaccard versus 50-61% for same-prompt reruns. Sampling more paraphrases per intent reduces the artifact in principle, and efficient multi-prompt evaluation methods exist in the academic literature, but the natural buyer-phrasing space is much larger than the benchmark-scale prompt sets those methods have been validated on, and far beyond what any commercial tracker issues per brand-intent combination. Prompt-by-prompt mention tracking is therefore structurally unstable as a unit of measurement; meaningful improvement likely requires a different unit rather than a larger prompt set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。