arXiv:2608.28576stat.MEcs.AI2026-08

用合成数据提升小样本推断,自动找最优数据量与权重组合。

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

  • 通过数据量和权重定义合成数据增强策略
  • 实验证明能达目标覆盖率并显著缩小置信区间
  • 适合小样本场景下需可靠推断的研究者

当真实数据稀缺时,合成数据可提升统计推断效果,但直接将合成样本视为真实数据会引入偏差,导致推断不可靠。本文提出一个跨相关任务群体的合成增强推断通用框架,以合成样本数量和权重刻画增强程度。核心是构建‘大小-权重前沿’:对每个权重,给出所有更小样本量均能达到目标任务边际覆盖的最大学样本量。该前沿基于历史任务估计,并在有限样本下对前沿上及以下所有配置同时保证覆盖性。在大语言模型生成回答以增强意见调查数据的实验中,该方法实现了目标覆盖性,并显著缩小了置信区间。

原文摘要 · Abstract (English)

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

合成数据推断优化小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。