arXiv:2505.15572cs.LGcs.AI2025-05被引 5

用强化学习让大模型更懂特定领域的数学公式生成

Bridging the Domain Gap in Equation Distillation with Reinforcement Feedback

  • 用下游数值表现作为奖励信号,直接优化生成策略
  • 在复杂数据分布下,方程准确率和鲁棒性显著提升
  • 适合需要物理可解释性的科研与工业建模场景

数据到方程(Data2Eqn)任务旨在发现将观测值映射到标签的可解释数学方程,提供物理洞察并广泛适用于学术与工业领域。遗传编程和传统深度学习方法在小规模任务数据集上存在搜索效率低、泛化能力差的问题。基础模型虽有潜力,但现有方法存在:1)在通用数据分布上预训练,难以适应领域特定任务;2)训练目标聚焦于词元级对齐,忽视数学语义,导致方程不准确。为此,我们提出一种基于强化学习的微调框架,通过下游数值适应度生成的奖励信号,直接优化预训练模型的生成策略。该方法使模型能够适应特定且复杂的数据分布,生成具有数学意义的方程。大量实验表明,该方法在复杂分布下显著提升了方程生成的准确性和鲁棒性。

原文摘要 · Abstract (English)

The data-to-equation (Data2Eqn) task aims to discover interpretable mathematical equations that map observed values to labels, offering physical insights and broad applicability across academic and industrial domains. Genetic programming and traditional deep learning-based approaches suffer from search inefficiency and poor generalization on small task-specific datasets. Foundation models showed promise in this area, but existing approaches suffer from: 1) They are pretrained on general-purpose data distributions, making them less effective for domain-specific tasks; and 2) their training objectives focus on token-level alignment, overlooking mathematical semantics, which can lead to inaccurate equations. To address these issues, we aim to enhance the domain adaptability of foundation models for Data2Eqn tasks. In this work, we propose a reinforcement learning-based finetuning framework that directly optimizes the generation policy of a pretrained model through reward signals derived from downstream numerical fitness. Our method allows the model to adapt to specific and complex data distributions and generate mathematically meaningful equations. Extensive experiments demonstrate that our approach improves both the accuracy and robustness of equation generation under complex distributions.

方程发现强化学习领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。