arXiv:2510.13665cs.LGcs.AI2025-10NeurIPS被引 1

提出可处理任意维度的神经网络,让物理模型更通用高效。

Axial Neural Networks for Dimension-Free Foundation Models

  • 用轴向结构共享参数,实现跨维度泛化
  • 在多维未见数据上表现优于传统模型
  • 适合需要跨尺度建模的物理仿真场景

AI 领域的基础模型显著提升了通用学习能力,支持零样本推理和上下文学习。然而,训练物理数据(如偏微分方程解)时面临维度不一的挑战:传统方法或固定最大维度,或为不同维度设独立编码器,导致效率低下。为此,我们提出一种维度无关的神经网络架构——轴向神经网络(XNN),受 Deep Sets 与图神经网络等参数共享结构启发。XNN 在保持计算效率的同时,能泛化至不同张量维度。我们将现有 PDE 基础模型转化为 XNN,评估三种训练场景:从头训练、多 PDE 预训练、单 PDE 微调。实验表明,XNN 性能与原模型相当,并在未见维度上展现出更优泛化能力,凸显多维预训练对基础模型的重要性。

原文摘要 · Abstract (English)

The advent of foundation models in AI has significantly advanced general-purpose learning, enabling remarkable capabilities in zero-shot inference and in-context learning. However, training such models on physics data, including solutions to partial differential equations (PDEs), poses a unique challenge due to varying dimensionalities across different systems. Traditional approaches either fix a maximum dimension or employ separate encoders for different dimensionalities, resulting in inefficiencies. To address this, we propose a dimension-agnostic neural network architecture, the Axial Neural Network (XNN), inspired by parameter-sharing structures such as Deep Sets and Graph Neural Networks. XNN generalizes across varying tensor dimensions while maintaining computational efficiency. We convert existing PDE foundation models into axial neural networks and evaluate their performance across three training scenarios: training from scratch, pretraining on multiple PDEs, and fine-tuning on a single PDE. Our experiments show that XNNs perform competitively with original models and exhibit superior generalization to unseen dimensions, highlighting the importance of multidimensional pretraining for foundation models.

基础模型偏微分方程维度泛化神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。