为物理科学图像设计高保真离散化方法,提升模拟精度与泛化能力。
Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Science
- 基于形状增益量化与本征正交分解,提出新离散化框架Phaedra。
- 在多类偏微分方程数据上重建误差降低,保留物理与谱空间特性。
- 对未知方程及真实气象观测数据具备强泛化能力,适合科学建模场景。
Token 是将高维数据转化为可高效学习、生成和泛化的序列的离散表示,已成为图像、视频生成及物理模拟的基础。然而,现有分词器针对真实视觉感知设计,未必适用于具有大动态范围、需保留物理与光谱特性的科学图像。本文评估多种图像分词器在刻画偏微分方程(PDE)物理与谱空间保真度方面的表现,发现其难以同时捕捉细节与精确量级。为此,我们提出受经典形状增益量化与本征正交分解启发的 Phaedra 框架。实验表明,Phaedra 在多个 PDE 数据集上持续提升重建质量。此外,其在三类递增复杂度任务中表现优异:已知方程不同条件下推理、未知方程推断,以及真实地球观测与气象数据的外分布泛化。
原文摘要 · Abstract (English)
Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. These have become foundational for image and video generation and, more recently, physical simulation. As existing tokenizers are designed for the explicit requirements of realistic visual perception of images, it is necessary to ask whether these approaches are optimal for scientific images, which exhibit a large dynamic range and require token embeddings to retain physical and spectral properties. In this work, we investigate the accuracy of a suite of image tokenizers across a range of metrics designed to measure the fidelity of PDE properties in both physical and spectral space. Based on the observation that these struggle to capture both fine details and precise magnitudes, we propose Phaedra, inspired by classical shape-gain quantization and proper orthogonal decomposition. We demonstrate that Phaedra consistently improves reconstruction across a range of PDE datasets. Additionally, our results show strong out-of-distribution generalization capabilities to three tasks of increasing complexity, namely known PDEs with different conditions, unknown PDEs, and real-world Earth observation and weather data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。