arXiv:2510.21977cs.AI2025-10ACL综述被引 4

用分布偏移对齐提升大模型模拟调查回答的准确性

Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions

  • 两阶段微调,学习分布变化而非直接拟合训练数据
  • 在5个公开数据集上显著优于其他方法,真实数据需求减少超50%
  • 适合需要低成本高效生成真实分布调查数据的研究者

大语言模型(LLMs)为模拟人类调查回答提供了低成本方案,但现有零样本方法受提示敏感性影响且准确率低,传统微调则过度拟合训练集分布,无法超越训练数据本身。为此,我们提出分布偏移对齐(DSA),一种两阶段微调方法,通过同时对齐输出分布及不同背景下的分布偏移,使模型学习分布变化规律而非直接拟合训练数据。实验表明,DSA在五个公开调查数据集上持续领先,可将所需真实数据量减少53.48%-69.12%,显著提升模拟精度与效率。

原文摘要 · Abstract (English)

Large language models (LLMs) offer a promising way to simulate human survey responses, potentially reducing the cost of large-scale data collection. However, existing zero-shot methods suffer from prompt sensitivity and low accuracy, while conventional fine-tuning approaches mostly fit the training set distributions and struggle to produce results more accurate than the training set itself, which deviates from the original goal of using LLMs to simulate survey responses. Building on this observation, we introduce Distribution Shift Alignment (DSA), a two-stage fine-tuning method that aligns both the output distributions and the distribution shifts across different backgrounds. By learning how these distributions change rather than fitting training data, DSA can provide results substantially closer to the true distribution than the training data. Empirically, DSA consistently outperforms other methods on five public survey datasets. We further conduct a comprehensive comparison covering accuracy, robustness, and data savings. DSA reduces the required real data by 53.48-69.12%, demonstrating its effectiveness and efficiency in survey simulation.

大模型调查模拟分布对齐数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。