arXiv:2509.20345stat.MEcs.LG2025-09被引 6

用合成数据提升小样本推断效率,质量差时自动切换回真实数据

General Synthetic-Powered Inference

  • 融合高质量合成数据与真实数据,动态调整推断策略
  • 误差率可控且随合成数据质量提升而下降
  • 兼容各类推断方法,适合数据稀缺场景

高质量合成数据的快速普及为统计推断带来了机遇与挑战。本文提出通用合成数据驱动推断框架(GESPI),可适配多种统计推断方法,通过结合合成与真实数据安全提升样本效率。该框架利用优质合成数据增强统计功效,当合成数据质量低时则自动退化为仅使用真实数据的标准方法。所提方法的误差率始终低于用户设定阈值,且无需对合成数据分布做任何假设;随着合成数据质量提高,误差率进一步降低。该灵活性使其可无缝集成到分位数预测、风险控制、假设检验及多重检验等流程中,且无需修改原有推断方法。我们在标签数据有限的复杂任务中验证了该方法的优势,包括AlphaFold蛋白质结构预测以及大型推理模型在复杂数学问题上的比较。

原文摘要 · Abstract (English)

The rapid proliferation of high-quality synthetic data -- generated by advanced AI models or collected as auxiliary data from related tasks -- presents both opportunities and challenges for statistical inference. This paper introduces a GEneral Synthetic-Powered Inference (GESPI) framework that wraps around a broad class of statistical inference procedures to safely enhance sample efficiency by combining synthetic and real data. Our framework leverages high-quality synthetic data to boost statistical power, yet adaptively defaults to the standard method using only real data when synthetic data are of low quality. The error rate of our method remains below a user-specified bound without any distributional assumptions on the synthetic data, and decreases as the quality of the synthetic data improves. This flexibility enables seamless integration with conformal prediction, risk control, hypothesis testing, and multiple testing procedures, all without modifying the base inference method. We demonstrate the benefits of our method on challenging tasks with limited labeled data, including AlphaFold protein structure prediction, and comparing large reasoning models on complex math problems.

统计推断合成数据小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。