用任务可交换性原则,让合成数据也能做靠谱的科学推断。
Valid Inference with Synthetic Data via Task Exchangeability
- 基于任务可交换性假设,建立合成数据下的有效推断方法
- 在民意调查和AI评估中验证了方法的统计有效性
- 适合需要快速迭代但又要求结果可信的研究者
合成数据在科学研究中的应用日益广泛,例如社会科学家使用大模型生成的“硅样本”进行预研,AI评估依赖大模型作为评判者,蛋白质组学则利用生成模型构建合成蛋白结构。这些进展带来新可能:研究者能提出更多问题、开展更多实验、加速发现。但同时存在根本风险:合成数据可能带有偏差、噪声或模型错配。本文提出一套基于统计原理的合成数据使用框架,具备可证明的有效性保证。核心思想是引入“任务可交换性”这一新条件——即研究者能识别出历史任务(真实数据可用),使得当前任务与历史任务在数学上可互换。我们发展了在任务可交换性下有效推断的方法,并拓展到超出该条件的情形。在公开民意调查使用硅样本及自动化评分器评估AI性能的场景中进行了验证。
原文摘要 · Abstract (English)
There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing for the use of LLM-generated "silicon samples" in pilot studies; AI evaluations increasingly rely on "LLM-as-a-judge" outputs; and proteomics research is accelerated by generative models that produce synthetic protein structures. These developments raise an intriguing possibility: synthetic data may help researchers ask more questions, run more studies, and accelerate discovery. But they also raise a fundamental concern: synthetic data can be biased, noisy, and misspecified. In this work, we propose statistical principles for using synthetic data in scientific research with provable validity guarantees. The key insight is a new technical condition that we call task exchangeability. Informally, this is a requirement that the researcher can identify historical tasks, for which real data is available, such that their current task of interest is exchangeable with the historical tasks in an appropriate mathematical sense. We develop methods for valid inference under task exchangeability, together with extensions that provide guarantees even beyond exchangeability. We demonstrate the framework on public opinion surveys with silicon samples and AI evaluation with autoraters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。