arXiv:2508.06635cs.LGcs.AI2025-08NeurIPS被引 8

用合成数据提升小样本研究的统计有效性,无需调参且理论可靠。

Valid Inference with Imperfect Synthetic Data

  • 基于广义矩估计设计新方法,自动融合真实与合成数据。
  • 合成数据与真实数据残差相关时,参数估计精度显著提升。
  • 适合社会科学研究者在数据稀缺时提升分析可靠性。

大语言模型生成的预测和合成数据正被越来越多地用于数据稀缺场景,如计算社会科学和人类受试者研究。尽管已有工作关注如何合理使用模型预测标签,但利用大语言模型生成全新合成样本(如模拟调查回答)并结合真实数据进行有效推断仍缺乏明确方法。本文提出一种基于广义矩估计的新估计器,实现无超参数调优的解决方案,并具备坚实的理论保证。有趣的是,我们发现合成数据与真实数据的矩残差之间存在相互预测关系时,能显著改善目标参数估计效果。通过多个计算社会科学任务的实证验证,该方法在有限样本下表现出显著性能提升。

原文摘要 · Abstract (English)

Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored the potential to use model-predicted labels for unlabeled data in a principled manner, there is increasing interest in using large language models to generate entirely new synthetic samples (e.g., synthetic simulations), such as in responses to surveys. However, it remains unclear by what means practitioners can combine such data with real data and yet produce statistically valid conclusions upon them. In this paper, we introduce a new estimator based on generalized method of moments, providing a hyperparameter-free solution with strong theoretical guarantees to address this challenge. Intriguingly, we find that interactions between the moment residuals of synthetic data and those of real data (i.e., when they are predictive of each other) can greatly improve estimates of the target parameter. We validate the finite-sample performance of our estimator across different tasks in computational social science applications, demonstrating large empirical gains.

合成数据统计推断小样本语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。