arXiv:2608.11937cs.LG2026-08

用教师模型生成长序列数据,让小模型高效学习偏微分方程求解。

Distillation of Foundation Models for Time-dependent PDEs

论文配图:Distillation of Foundation Models for Time-dependent PDEs
图 1 · 摘自论文原文
  • 通过教师模型滚动生成合成轨迹扩充数据
  • 学生模型参数量减少数个数量级,推理速度提升超10倍
  • 适合需要快速求解的物理模拟场景

针对时间依赖型偏微分方程(PDE)的基础模型在大量物理系统上训练,经少量目标域轨迹微调后可在低数据条件下实现高精度。但其规模大、计算成本高,难以作为数值求解器的快速代理。本文提出教师滚动扩展(TREX)知识蒸馏框架,从微调后的教师模型出发,通过教师滚动生成长时序合成轨迹(可选周期性噪声注入),在无需初始条件分布先验的情况下采样教师诱导的滚动分布。该过程使学生模型接触到长时序状态及自回归预测中的局部恢复行为,并可融入任务特定归纳偏置(如等变性)。在多个PDE基准上评估显示,所得学生模型在参数量减少数个数量级的同时,精度可匹配或超越教师模型,推理速度提升超过一个数量级。

原文摘要 · Abstract (English)

Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few trajectories from a target domain, they can achieve strong accuracy in low-data regimes. However, these models are typically large and computationally intensive, limiting their usefulness as fast surrogates for numerical solvers. We propose Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student. Starting from a fine-tuned teacher, TREX augments limited downstream data by generating long synthetic trajectories through teacher rollouts, optionally with periodic noise injection. This procedure samples from the teacher-induced rollout distribution without requiring explicit knowledge of the initial-condition distribution, while exposing the student to long-horizon states and local recovery behavior around states encountered during autoregressive prediction. The student can further incorporate task-specific inductive biases, such as equivariance, that the teacher does not necessarily enforce. We evaluate TREX on multiple PDE benchmarks. The resulting students can match or surpass the teacher's accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.

PDE求解知识蒸馏高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。