arXiv:2603.13227cs.LGcs.CV2026-03被引 5

对比自监督方法在物理系统表征学习中的表现,发现隐空间模型更优。

Representation Learning for Spatiotemporal Physical Systems

  • 用下游科学任务评估表征质量,而非仅预测下一帧。
  • 隐空间学习方法(如JEPAs)在参数估计上优于像素级预测模型。
  • 通用自监督方法有时比专门物理建模方法更有效,值得科研人员关注。

针对时空物理系统的机器学习方法多聚焦于下一帧预测,旨在学习系统随时间演化的精确模拟器。然而,这类模拟器训练成本高,且在自回归推理中易出现误差累积。本文另辟蹊径,关注预测下一帧之后的下游科学任务,如系统控制参数的估计。这些任务的准确率能定量反映模型表征的物理意义。我们评估了通用自监督方法在学习对下游科学任务有用的物理基础表征方面的效果。令人意外的是,并非所有专为物理建模设计的方法都优于通用自监督方法;相反,在隐空间中学习的方法(如联合嵌入预测架构,JEPAs)在参数估计任务中表现更佳。代码已开源:https://github.com/helenqu/physical-representation-learning。

原文摘要 · Abstract (English)

Machine learning approaches to spatiotemporal physical systems have primarily focused on next-frame prediction, with the goal of learning an accurate emulator for the system's evolution in time. However, these emulators are computationally expensive to train and are subject to performance pitfalls, such as compounding errors during autoregressive rollout. In this work, we take a different perspective and look at scientific tasks further downstream of predicting the next frame, such as estimation of a system's governing physical parameters. Accuracy on these tasks offers a uniquely quantifiable glimpse into the physical relevance of the representations of these models. We evaluate the effectiveness of general-purpose self-supervised methods in learning physics-grounded representations that are useful for downstream scientific tasks. Surprisingly, we find that not all methods designed for physical modeling outperform generic self-supervised learning methods on these tasks, and methods that learn in the latent space (e.g., joint embedding predictive architectures, or JEPAs) outperform those optimizing pixel-level prediction objectives. Code is available at https://github.com/helenqu/physical-representation-learning.

表征学习物理系统自监督隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。