用大模型模拟用户可替代真实实验,加速算法评估
LLM Personas as a Substitute for Field Experiments in Method Benchmarking
- 用大模型角色模拟人类,在仅观察整体结果时可替代真实用户
- 只需足够多独立角色评估,就能可靠区分不同算法效果
- 适合快速测试算法、无需真实数据的场景,如政策模拟
真实社会系统中的方法评估常依赖成本高、周期长的实地实验(A/B测试),阻碍方法迭代。本文提出用大模型生成的角色模拟替代真实用户,但关键问题是:这种替换是否保持评估接口的一致性?我们证明:当方法仅能观测整体结果(仅聚合观测),且评估仅关注提交成果而非方法来源(方法无关评估)时,将真人替换为角色等价于更换评估人群(如纽约换雅加达)。进一步地,我们引入信息论意义上的聚合通道可区分性,表明让角色评估与实地实验一样具备决策价值,本质上是样本量问题,并给出在特定分辨力下区分不同方法所需的最小独立角色评估数的明确边界。
原文摘要 · Abstract (English)
Field experiments (A/B tests) are often the most credible benchmark for methods (algorithms) in societal systems, but their cost and latency bottleneck rapid methodological progress. LLM-based persona simulation offers a cheap synthetic alternative, yet it is unclear whether replacing humans with personas preserves the benchmark interface that adaptive methods optimize against. We prove an if-and-only-if characterization: when (i) methods observe only the aggregate outcome (aggregate-only observation) and (ii) evaluation depends only on the submitted artifact and not on the method's identity or provenance (method-blind evaluation), swapping humans for personas is just panel change from the method's point of view, indistinguishable from changing the evaluation population (e.g., New York to Jakarta). Furthermore, we move from validity to usefulness: we define an information-theoretic discriminability of the induced aggregate channel and show that making persona benchmarking as decision-relevant as a field experiment is fundamentally a sample-size question, yielding explicit bounds on the number of independent persona evaluations required to reliably distinguish meaningfully different methods at a chosen resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。