arXiv:2604.17267cs.AIstat.AP2026-04

用大模型生成问卷回复,智能分配真人审核资源,省钱提效。

Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys

  • 根据题目难易度动态分配真人标注量,越不准的题给越多真人
  • 实测比传统方法误差降低10.5%~11.4%,无需试填数据
  • 新任务也能预测难度,适合快速部署的市场调研场景

大型语言模型可低成本生成合成问卷回答,但其准确性在不同问题间波动不定。本文研究在固定人力预算下,如何合理分配真人回答样本以优化估计任务。框架包含三部分:首先,基于预测驱动推断,定义了每个问题特有的校正难度,决定了估计方差随真人样本量增长的收敛速度;其次,推导出闭式最优分配规则,将更多真人标注分配给大模型表现最差的任务;第三,由于校正难度依赖未观测的人类回答,提出一种基于历史数据训练的元学习方法,可预测全新任务的难度而无需预调查数据。该框架扩展至一般M-估计,涵盖回归系数和联合分析中的多项对数几率效用值。在两个跨领域、多题型、多模型的数据集上验证,所提方法实现了理论效率提升的61%-79%,在无任何目标调查预调查数据条件下,分别实现11.4%和10.5%的均方误差下降。

原文摘要 · Abstract (English)

Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LLM predictions are available for every task. Our framework combines three components. First, building on Prediction-Powered Inference, we characterize a question-specific rectification difficulty that governs how quickly the estimator's variance decreases with human sample size. Second, we derive a closed-form optimal allocation rule that directs more human labels to tasks where the LLM is least reliable. Third, since rectification difficulty depends on unobserved human responses for new surveys, we propose a meta-learning approach, trained on historical data, that predicts it for entirely new tasks without pilot data. The framework extends to general M-estimation, covering regression coefficients and multinomial logit partworths for conjoint analysis. We validate the framework on two datasets spanning different domains, question types, and LLMs, showing that our approach captures 61-79% of the theoretically attainable efficiency gains, achieving 11.4% and 10.5% MSE reductions without requiring any pilot human data for the target survey.

问卷设计大模型应用资源分配效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。