用强化学习让大模型自适应生成符合物理规律的简洁方程
LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
- 通过双重约束机制,让大模型在生成方程时兼顾物理合理性和结构简洁性
- 在标准测试中表现超越现有方法,成功发现新型湍流模型
- 小模型经训练可超越闭源大模型,降低科学发现门槛
符号回归旨在从观测数据中提炼数学方程。近期方法利用大语言模型(LLMs)生成方程假设,借助其丰富的预训练科学先验。然而,现有框架多将LLM视为静态生成器,仅依赖提示词引导探索,无法根据搜索反馈更新内部表征,常导致物理不一致或数学冗余表达。本文提出PiT-PO(Physics-informed Token-regularized Policy Optimization)统一框架,通过强化学习使LLM演变为自适应生成器。核心是双约束机制,严格保证层级物理有效性,同时施加细粒度的标记级惩罚以抑制冗余结构。因此,PiT-PO使模型生成兼具科学一致性与结构简洁性的方程。实验证明,PiT-PO在标准基准上达到当前最优性能,并成功为复杂流体力学问题发现新型湍流模型。我们还展示,经过训练的小规模模型可超越闭源大模型,实现高性能科学发现的普惠化。
原文摘要 · Abstract (English)
Symbolic regression aims to distill mathematical equations from observational data. Recent approaches have successfully leveraged Large Language Models (LLMs) to generate equation hypotheses, capitalizing on their vast pre-trained scientific priors. However, existing frameworks predominantly treat the LLM as a static generator, relying on prompt-level guidance to steer exploration. This paradigm fails to update the model's internal representations based on search feedback, often yielding physically inconsistent or mathematically redundant expressions. In this work, we propose PiT-PO (Physics-informed Token-regularized Policy Optimization), a unified framework that evolves the LLM into an adaptive generator via reinforcement learning. Central to PiT-PO is a dual-constraint mechanism that rigorously enforces hierarchical physical validity while simultaneously applying fine-grained, token-level penalties to suppress redundant structures. Consequently, PiT-PO aligns LLM to produce equations that are both scientifically consistent and structurally parsimonious. Empirically, PiT-PO achieves state-of-the-art performance on standard benchmarks and successfully discovers novel turbulence models for challenging fluid dynamics problems. We also demonstrate that PiT-PO empowers small-scale models to outperform closed-source giants, democratizing access to high-performance scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。