arXiv:2509.13136cs.LG2025-09被引 1

用扩散模型生成可解释的数学方程,提升科学发现效率。

Discovering Mathematical Equations with Diffusion Language Model

  • 基于连续扩散语言模型,将符号映射到隐空间进行方程建模。
  • 在标准数据集上表现媲美顶尖自回归方法,生成更多样且可读方程。
  • 适合需要自动发现物理规律的研究者,尤其擅长复杂方程搜索。

从观测数据中发现有效且有意义的数学方程,在科学发现中至关重要。这一任务即符号回归,因搜索空间巨大且需权衡精度与复杂度而极具挑战。本文提出 DiffuSR,一个基于连续状态扩散语言模型的预训练框架,用于符号回归。DiffuSR 在扩散过程中引入可训练嵌入层,将离散数学符号映射至连续隐空间,有效建模方程分布。通过迭代去噪,将初始噪声序列转化为符号方程,数值数据通过交叉注意力机制注入引导。我们还设计了高效的推理策略,将对数先验注入遗传编程以提升生成准确性。在标准符号回归基准上的实验表明,DiffuSR 性能媲美最先进自回归方法,并生成更具可解释性和多样性的数学表达式。

原文摘要 · Abstract (English)

Discovering valid and meaningful mathematical equations from observed data plays a crucial role in scientific discovery. While this task, symbolic regression, remains challenging due to the vast search space and the trade-off between accuracy and complexity. In this paper, we introduce DiffuSR, a pre-training framework for symbolic regression built upon a continuous-state diffusion language model. DiffuSR employs a trainable embedding layer within the diffusion process to map discrete mathematical symbols into a continuous latent space, modeling equation distributions effectively. Through iterative denoising, DiffuSR converts an initial noisy sequence into a symbolic equation, guided by numerical data injected via a cross-attention mechanism. We also design an effective inference strategy to enhance the accuracy of the diffusion-based equation generator, which injects logit priors into genetic programming. Experimental results on standard symbolic regression benchmarks demonstrate that DiffuSR achieves competitive performance with state-of-the-art autoregressive methods and generates more interpretable and diverse mathematical expressions.

符号回归扩散模型数学发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。