通过特征互换识别假数据,让模型学会物理一致性规律。
Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations

- 用其他实体的真实特征替换部分属性,训练模型识别被调换的数据
- 在21个环境变量上验证,显著提升气候分类等7项下游任务表现
- 适合做科学建模的自监督学习,尤其关注物理规律的模型
科学数据中的实体特征通常受物理定律共同约束,但现有自监督学习目标大多忽视这种物理一致性。我们提出imposter,一种判别式预训练任务:将某实体的部分特征替换为另一实体的真实观测值,训练编码器识别被替换的特征。由于每个替换值本身都合理,该任务只能通过学习跨特征的物理依赖关系来解决。我们在全球ERA5-Land再分析数据集上使用21个环境变量进行评估,并在气候分类、碳通量估计和径流预测等七项下游任务中测试所学表征。据我们所知,这是首次在统一架构与预训练预算下对地表建模自监督目标的系统比较。结果表明,最有效的预训练任务取决于下游任务类型,而非单一目标优越性;且imposter能与现有目标互补。这表明物理一致性是科学基础模型的重要自监督信号。
原文摘要 · Abstract (English)
Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to identify the swapped features. Because every donated value is individually plausible, the task can only be solved by learning cross-feature physical dependencies. We evaluate the proposed objectives on global ERA5-Land reanalysis data using 21 environmental variables and assess the learned representations on seven downstream tasks spanning climate classification, carbon flux estimation, and streamflow prediction. Our study includes, to our knowledge, the first systematic comparison of self-supervised objectives for land-surface modeling under a shared architecture and pre-training budget. We find that the most effective pretext task depends on the downstream task family rather than any single objective's superiority, and that imposter provides complementary information when combined with existing SSL objectives. These results suggest that physical coherence is a valuable new source of self-supervision for scientific foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。