arXiv:2506.04282cs.LG2025-06被引 21

用数据与反思双轮驱动,提升大模型发现科学方程的能力

DrSR: LLM based Scientific Equation Discovery with Dual Reasoning from Data and Experience

  • 结合数据特征分析与生成反馈,形成闭环推理
  • 在多学科数据上有效方程率显著提升,优于传统与现有方法
  • 适合需要可解释方程的科研人员快速探索复杂规律

符号回归是通过数据发现可解释数学表达式的基石工具,在科学与工程领域应用广泛。近年来,大语言模型(LLM)凭借内置科学先验和推理能力,在该任务中表现优异,超越传统方法。然而,现有基于LLM的方法(如LLM-SR)过度依赖内部先验,缺乏对数据的显式理解及生成过程中的系统性反思。为此,我们提出DrSR(Dual Reasoning Symbolic Regression),通过融合数据驱动洞察与反思学习,增强模型的鲁棒性与发现能力。具体而言,DrSR引导LLM分析数据中的结构关系(如单调性、非线性、相关性),生成结构化描述;同时监控方程性能,建立反馈回路以优化后续生成。通过将数据理解与生成反思集成于闭环中,DrSR实现对符号表达空间更高效的探索。在物理、化学、生物及材料科学等跨学科数据集上的实验表明,DrSR显著提升有效方程率,在准确性、泛化能力与搜索效率方面持续优于经典与近期的基于LLM方法,展现出其在科学方程发现中的巨大潜力。

原文摘要 · Abstract (English)

Symbolic regression is a fundamental tool for discovering interpretable mathematical expressions from data, with broad applications across scientific and engineering domains. Recently, large language models (LLMs) have demonstrated strong performance in this task, leveraging embedded scientific priors and reasoning capabilities to surpass traditional methods. However, existing LLM-based approaches, such as LLM-SR, often over-rely on internal priors, lacking explicit data understanding and systematic reflection during equation generation. To address these limitations, we propose DrSR (Dual Reasoning Symbolic Regression), a framework that combines data-driven insight with reflective learning to enhance both robustness and discovery capability. Specifically, DrSR guides LLMs to analyze structural relationships (e.g., monotonicity, nonlinearity, and correlation) within the data to generate structured descriptions. Simultaneously, it monitors equation performance and establishes a feedback loop to refine subsequent generations. By integrating data understanding and generation reflection in a closed loop, DrSR enables more efficient exploration of the symbolic expression space. Experiments across interdisciplinary datasets in physics, chemistry, biology, and materials science demonstrate that DrSR substantially improves the valid equation rate and consistently outperforms both classical and recent LLM-based methods in terms of accuracy, generalization, and search efficiency. These results underscore its potential for scientific equation discovery.

符号回归科学发现大模型方程生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。