用推理时计算提升微分方程预测精度,少训练、小模型也能准。
Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
- 推理时动态调用计算资源,通过奖励模型评估时空一致性
- 在PDEGym上比传统自回归方法误差降低37%,尤其擅长分布外场景
- 适合追求高精度但算力有限的物理模拟与工程应用
偏微分方程(PDE)是现代计算科学与工程的基础,但求解过程计算成本高。尽管PDE基础模型在模拟复杂时空现象方面展现出潜力,现有模型受限于预训练数据,自回归推演性能差,尤其在分布外(OOD)情况下表现不佳,且对算力和训练数据需求大。受大语言模型“思考”策略启发,本文首次提出针对PDE的测试时计算(TTC)策略,在推理阶段利用计算资源提升预测精度,仅需少量训练样本和较小模型即可实现。通过两类基于随机模型的奖励模型评估时空一致性,该方法在PDEGym基准的可压缩欧拉方程仿真中显著优于标准自回归推理。TTC框架为构建更高级的推理算法与强化学习驱动的PDE建模奠定基础,可能变革物理与工程中的计算流程。
原文摘要 · Abstract (English)
Partial Differential Equations (PDEs) are the bedrock for modern computational sciences and engineering, and inherently computationally expensive. While PDE foundation models have shown much promise for simulating such complex spatio-temporal phenomena, existing models remain constrained by the pretraining datasets and struggle with auto-regressive rollout performance, especially in out-of-distribution (OOD) cases. Furthermore, they have significant compute and training data requirements which hamper their use in many critical applications. Inspired by recent advances in ``thinking" strategies used in large language models (LLMs), we introduce the first test-time computing (TTC) strategy for PDEs that utilizes computational resources during inference to achieve more accurate predictions with fewer training samples and smaller models. We accomplish this with two types of reward models that evaluate predictions of a stochastic based model for spatio-temporal consistency. We demonstrate this method on compressible Euler-equation simulations from the PDEGym benchmark and show that TTC captures improved predictions relative to standard non-adaptive auto-regressive inference. This TTC framework marks a foundational step towards more advanced reasoning algorithms or PDE modeling, inluding building reinforcement-learning-based approaches, potentially transforming computational workflows in physics and engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。