arXiv:2506.17774cs.LG2025-06被引 27

首个物理仿真大模型,用离散化序列生成物理过程。

PhysiX: A Foundation Model for Physics Simulations

  • 用离散令牌编码多尺度物理过程,自回归建模序列演化
  • 45亿参数模型在真实数据集上超越专用模型与现有最佳方法
  • 首次证明自然视频知识可迁移至物理仿真,支持多任务协同学习

基础模型已在视频、图像和语言领域取得显著成功。通过扩大参数量和训练数据规模,这些模型获得了可泛化的世界知识,并常优于特定任务模型。然而,这一进展尚未延伸至物理仿真领域。主要瓶颈在于数据稀缺:尽管互联网上有数百万张图片、视频和文本资源,但最大的物理仿真数据集仅包含数万样本。这种数据限制阻碍了大模型的应用,过拟合成为主要问题。因此,物理应用通常依赖小模型,其因上下文理解能力有限而难以进行长程预测。此外,与图像、视频或文本的固定粒度不同,物理数据集在尺度上差异极大,进一步加剧了多任务训练的难度。我们提出 PhysiX,首个大规模物理仿真基础模型。PhysiX 是一个 45 亿参数的自回归生成模型,采用离散分词器将不同尺度的物理过程编码为离散令牌序列,并使用自回归下一令牌预测目标在令牌空间中建模这些过程。为缓解离散化过程中的舍入误差,PhysiX 引入了专用精炼模块。大量实验表明,PhysiX 有效克服了数据瓶颈,在与基准模型相当设置下表现优于任务专用基线,并在 The Well 基准测试中达到当前绝对最优水平。结果表明,自然视频中学到的知识可成功迁移到物理仿真中,且跨多样化仿真任务的联合训练能实现协同学习。

原文摘要 · Abstract (English)

Foundation models have achieved remarkable success across video, image, and language domains. By scaling up the number of parameters and training datasets, these models acquire generalizable world knowledge and often surpass task-specific approaches. However, such progress has yet to extend to the domain of physics simulation. A primary bottleneck is data scarcity: while millions of images, videos, and textual resources are readily available on the internet, the largest physics simulation datasets contain only tens of thousands of samples. This data limitation hinders the use of large models, as overfitting becomes a major concern. As a result, physics applications typically rely on small models, which struggle with long-range prediction due to limited context understanding. Additionally, unlike images, videos, or text-which typically exhibit fixed granularity-physics datasets often vary drastically in scale, amplifying the challenges of scaling up multitask training. We introduce PhysiX, the first large-scale foundation model for physics simulation. PhysiX is a 4.5B parameter autoregressive generative model. It uses a discrete tokenizer to encode physical processes at different scales into a sequence of discrete tokens, and employs an autoregressive next-token prediction objective to model such processes in the token space. To mitigate the rounding error in the discretization process, PhysiX incorporates a specialized refinement module. Through extensive experiments, we show that PhysiX effectively addresses the data bottleneck, outperforming task-specific baselines under comparable settings as well as the previous absolute state-of-the-art approaches on The Well benchmark. Our results indicate that knowledge learned from natural videos can be successfully transferred to physics simulation, and that joint training across diverse simulation tasks enables synergistic learning.

物理仿真基础模型自回归离散化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。