用扩散模型生成数据,提升小样本下自旋系统参数推断效果
Data augmentation using diffusion models to enhance inverse Ising inference
- 用扩散模型生成符合物理规律的新数据,扩充小样本集
- 在合成数据和神经活动缺失值重建中均显著改善推断精度
- 适合物理建模与小样本数据科学问题的研究者
从观测配置中识别模型参数是数据科学中的基础挑战,尤其在数据有限时更为突出。近年来,扩散模型作为生成式机器学习的新范式,能够生成高度贴近真实数据的新样本,其通过学习模型概率梯度,避免了对所有可能配置求解分区函数的复杂计算。本文探讨了扩散模型能否通过数据增强提升小样本下的参数推断能力。研究通过一个合成的逆伊辛推断任务和一个真实的神经活动数据缺失值重建应用验证了该方法的有效性。结果表明,利用扩散模型生成的数据可显著改善参数估计性能,为物理相关问题的数据增强提供了概念验证,拓展了数据科学的新路径。
原文摘要 · Abstract (English)
Identifying model parameters from observed configurations poses a fundamental challenge in data science, especially with limited data. Recently, diffusion models have emerged as a novel paradigm in generative machine learning, capable of producing new samples that closely mimic observed data. These models learn the gradient of model probabilities, bypassing the need for cumbersome calculations of partition functions across all possible configurations. We explore whether diffusion models can enhance parameter inference by augmenting small datasets. Our findings demonstrate this potential through a synthetic task involving inverse Ising inference and a real-world application of reconstructing missing values in neural activity data. This study serves as a proof-of-concept for using diffusion models for data augmentation in physics-related problems, thereby opening new avenues in data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。