arXiv:2605.27463stat.MEcs.AI2026-05被引 1

LLM生成调研易受提示词微调影响,新方法确保统计有效性

When prompt perturbations break your A/B test: A valid statistical test for generative surveying

论文配图:When prompt perturbations break your A/B test: A valid statistical test for generative surveying
图 1 · 摘自论文原文
  • 提出基于置换检验的新统计方法,应对提示扰动带来的偏差
  • 实证显示标准检验在多数情况下会误判结果,存在显著假阳性风险
  • 适用于依赖大模型生成反馈的市场调研、产品测试等场景

生成式调研——利用基于大语言模型的虚拟人物对信息进行反馈——已成为低成本、可扩展的传统市场研究替代方案。然而,大语言模型对提示词设计的细微变化极为敏感,调研结论可能依赖于随意的措辞选择。为控制这种敏感性,分析中需包含语义等价的提示扰动。本文表明,在包含真实扰动结构的统计模型下,标准假设检验(如符号检验和威尔科克斯符号秩检验)是无效的。我们提出一种在该模型下有效的置换检验,并形式化刻画了标准检验失效的条件。通过一个简单生成式调研问题的应用,我们估计了相关参数,评估了置换检验在现实条件下的检验力,并提供了跨人物、扰动与重复次数的预算分配建议。最后,我们发现即使在同一模型族内,估计效应的大小与方向也高度依赖于具体模型选择。

原文摘要 · Abstract (English)

Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to traditional market research. However, LLMs are sensitive to small variations in prompt design and conclusions drawn from generative surveys may depend on arbitrary phrasing choices. Controlling for this sensitivity requires including semantically equivalent perturbations in the analysis. In this paper, we show that standard hypothesis tests, including the sign test and Wilcoxon signed-rank test, are invalid under a statistical model for generative surveying that includes realistic perturbation structure. We propose a permutation test that is valid under this model and formally characterize the conditions under which standard tests fail. Applying our framework to a simple generative surveying problem, we estimate relevant parameters, characterize the power of the permutation test under realistic conditions, and provide practical guidance on budget allocation across personas, perturbations, and replicates. Finally, we show that both the magnitude and direction of the estimated effect are sensitive to the choice of model, even within the same model family.

大模型调研统计验证提示敏感性置换检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。