用扩散模型生成符合物理规律的合成数据,提升稀疏观测下符号回归的准确性。
Data Enrichment for Symbolic Regression Using Diffusion Models

- 基于变分自编码器与物理引导扩散模型生成合成数据
- 在热传导等场景中显著提升稀疏条件下的方程恢复率
- 无需领域知识即可实现高质量数据增强,适合科研与工程应用
符号回归(SR)可通过观测数据发现可解释的物理方程,但当时空测量数据稀疏、噪声大或不完整时,其可靠性急剧下降。数据增强(DE)可缓解此问题,但若新增样本破坏系统物理结构,反而会误导方程发现。为此,本文提出一种物理引导的潜在扩散框架用于数据增强,结合变分自编码器、条件潜在扩散模型和物理信息残差校正器,生成符合控制关系的合成场以补全稀疏观测。在热传导、不可压缩纳维-斯托克斯流及单质量牛顿引力势等场景中,使用GPLearn、DEAP和PySR作为下游符号回归工具,结果表明:经物理校正的数据增强能持续提升各类动力学系统与符号回归模型在稀疏条件下的方程恢复性能。这证明生成式数据增强可在不依赖额外领域知识的前提下,有效强化方程发现能力。
原文摘要 · Abstract (English)
Symbolic regression (SR) offers a route to scientific discovery by converting observations into interpretable governing equations. However, despite its promise, its reliability degrades sharply when spatiotemporal measurements are sparse, noisy, or physically incomplete, as commonly occurring in practice. Data enrichment (DE) has been shown to be able to mitigate this limitation, yet additional samples can mislead equation discovery unless they preserve the physical structure of the target system. Such implication of DE requires narrow domain expertise as well as technical fluidity, highly limiting its practical usefulness. In this study, we introduce a physics-guided latent diffusion framework for DE for down the line SR models. The proposed framework combines a variational autoencoder, a conditional latent diffusion model, and a physics-informed residual corrector to complete sparse observations with synthetic fields constrained by governing relations. We evaluate the approach on heat conduction, incompressible Navier-Stokes flow, and a moving single-mass Newtonian gravitational potential, using GPLearn, DEAP, and PySR as downstream SR backends. Our results reveal that physics-corrected enrichment consistently improves recovery in sparse regimes across physical dynamics and SR models. These results show that generative enrichment can strengthen equation discovery without additional domain expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。