arXiv:2510.11408cs.CL2025-10ACL综述被引 15

用大模型生成问卷回复,再用校正方法大幅降低偏差,提升效率。

Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification

  • 用大模型合成回答,再通过校正减少偏差
  • 结合方法使偏差低于5%,样本量提升14倍
  • 预算有限时,多数人力应投入校正而非微调

问卷调查能提供公众意见与行为的宝贵见解,但执行成本高、耗时长。大语言模型(LLMs)被提出作为低成本、可扩展的人类受访者替代方案,但其输出常存在偏差,导致估计无效。本文研究基于LLM生成调查回答的合成方法与用于消除人口估计偏差的校正方法之间的协同作用,并探讨在两者间如何最优分配有限的人类数据。基于两项包含营养、政治和经济问题的面板调查,我们发现仅使用合成方法会引入24%-86%的显著偏差;而结合合成与校正后,偏差降至5%以下,有效样本量最高提升14%。结果表明,在固定预算下,将大部分人类数据用于校正而非全量用于微调,能实现更有效的估计。

原文摘要 · Abstract (English)

Surveys provide valuable insights into public opinion and behavior, but their execution is costly and slow. Large language models (LLMs) have been proposed as a scalable, low-cost substitute for human respondents, but their outputs are often biased and yield invalid estimates. We study the interplay between synthesis methods that use LLMs to generate survey responses and rectification methods that debias population estimates, and explore how human responses are best allocated between them. Using two panel surveys with questions on nutrition, politics, and economics, we find that synthesis alone introduces substantial bias (24-86%), whereas combining it with rectification reduces bias below 5% and increases effective sample size by up to 14%. Overall, we challenge the common practice of using all human responses for fine-tuning, showing that under a fixed budget, allocating most to rectification results in far more effective estimation.

大模型问卷模拟偏差校正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。