arXiv:2510.05278cs.LGcs.CL2025-10

让更流行的解码器模型也能高效处理偏微分方程模拟任务。

Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs

  • 提出两种模仿双向性的新方法,提升解码器模型性能。
  • 在时间依赖的PDE模拟任务中,性能接近编码器模型。
  • 适合希望用大模型做科学计算的研究者参考。

尽管大语言模型主要应用于自然语言任务,但其在跨模态适配方面也展现出巨大潜力,例如用于科学机器学习任务。现有大多数跨模态适配方法聚焦于编码器仅有的Transformer架构,而近年来解码器仅有的架构在语言任务中更为流行且训练规模更大。这引发了关于模型架构如何影响跨模态适配的问题,以及是否可利用解码器模型的成功。本文系统比较了编码器仅有和解码器仅有语言模型在基于偏微分方程(PDEs)的时间依赖模拟任务中的跨模态适配表现。结果发现,当直接应用现有方法时,解码器仅有模型远逊于编码器仅有模型;与其它领域不同,扩大解码器仅有模型规模也无帮助。为此,我们提出两种新方法:并行翻转(Parallel Flipping)和序列加倍(Sequence Doubling),均能显著提升所有任务和跨模态适配方法下解码器仅有模型的性能,缩小与编码器仅有模型的差距。希望本研究拓展跨模态适配任务可用模型范围,推动科学机器学习发展。

原文摘要 · Abstract (English)

While large language models are primarily used on natural language tasks, they have also shown great promise when adapted to new modalities, e.g., for scientific machine learning tasks. Most proposed approaches for such cross-modal adaptation of language models focus on encoder-only transformer model architectures, despite decoder-only architectures being far more popular for language tasks in recent years, and being trained at much larger scales. This raises the question of how model architecture affects cross-modal adaptation approaches, and whether we can leverage the success of decoder-only models. In this paper, we systematically compare encoder-only and decoder-only language models on cross-modal adaptation for time-dependent simulation tasks based on partial differential equations (PDEs). We find that decoder-only models are far worse than encoder-only models, when existing approaches are applied unmodified. In contrast to several other domains, scaling decoder-only models also does not help. To enhance the performance of decoder-only models in this context, we introduce two novel approaches that mimic bidirectionality, Parallel Flipping and Sequence Doubling. Both our methods improve overall performance using decoder-only models for all tasks and all cross-modal adaptation methods, closing the gap to encoder-only model performance. We hope that our findings broaden the spectrum of models used on cross-modal adaptation tasks to further scientific machine learning.

偏微分方程跨模态适配解码器模型科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。