让模型自动发现物理场的多尺度结构,无需标签或预设规则。
ScaleAware-JEPA: Latent Representation for Discovery in Multiscale Physical Fields

- 用约束扩散分解提取连续场的多尺度分量,生成隐式坐标。
- 预测任务基于各尺度分量的扩散特性,而非固定图像块。
- 在湍流、星际气体等数据上重建出清晰结构图谱,适合科学发现。
连续物理场是科学调查中大量存在的数据类型,其多尺度结构对发现至关重要,但通常缺乏先验坐标。标准自监督方法在固定图像坐标下定义上下文与目标,导致预测任务与跨尺度连续结构不匹配。本文提出ScaleAware-JEPA框架,通过约束扩散分解(CDD)将每个场分解为像素对齐的尺度分量,并提供定义掩码几何的尺度坐标。由此构建的JEPA目标使用与各分量扩散尺度相关的上下文区域进行隐藏结构预测,而非依赖任意补丁大小。在磁流体湍流、星际分子气体和城市夜间灯光结构等数据上,学习到的几何结构能回溯到连贯形态,形成无需标签或预设分割规则的密集结构图谱。通过将隐式预测绑定于场的尺度层级,ScaleAware-JEPA实现了复杂物理模式的可观察性,即便相关结构尚未被指定。代码已开源:https://github.com/gxli/SA-JEPA。
原文摘要 · Abstract (English)
Continuous physical fields represent a large fraction of data under scientific investigation. Their multiscale structures are central to discovery, yet useful coordinates are not known in advance. Standard self-supervised methods define context and targets in fixed image coordinates, posing a predictive task misaligned with fields organized across a continuous scale hierarchy. We introduce ScaleAware-JEPA, a framework that constructs dense, label-free latent coordinates for continuous scalar fields. Constrained Diffusion Decomposition (CDD) separates each field into pixel-registered scale components and provides the scale coordinates that define the masking geometry. The resulting JEPA objective predicts hidden structure with a context footprint tied to the diffusion scale of each component rather than to an arbitrary patch size. Across MHD turbulence, interstellar molecular gas and urban nighttime-light structure, the learned geometry maps back to coherent morphology, forming dense structural atlases without labels or predefined segmentation rules. By tying latent prediction to the scale hierarchy of a field, ScaleAware-JEPA constructs latent coordinates through which complex physical patterns can be inspected before their relevant structures have been prescribed. Code is available at https://github.com/gxli/SA-JEPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。