用少量数据对齐分布外样本特征,提升大规模观测的模型泛化能力。
Towards Mitigating Systematics in Large-Scale Surveys via Few-Shot Optimal Transport-Based Feature Alignment
- 基于最优传输的特征对齐损失,实现预训练模型在分布外样本上的特征匹配。
- 在中性氢大尺度巡天数据上验证,仅用少量样本即显著提升模型性能。
- 适用于系统误差未知、数据稀缺的真实天文巡天场景,适合天体物理研究者。
系统误差会污染观测数据,导致其分布偏离理论模拟信号,给使用预训练模型标注此类观测带来挑战。由于系统误差常难以理解且难建模,直接完全消除往往不可行。为此,我们提出一种新方法:通过优化从预训练分布内(ID)模型提取的表征之间的特征对齐损失,实现分布内与分布外(OOD)样本特征的对齐。我们在MNIST数据集上实验验证了多种对齐损失的有效性,包括均方误差和最优传输,并进一步将其应用于中性氢的大尺度分布图。结果表明,在分布内与分布外样本之间缺乏对称性的情况下,最优传输仍能有效对齐特征,即使在数据量有限时也表现良好,模拟了从大规模巡天中提取信息的真实条件。代码已公开于https://github.com/sultan-hassan/feature-alignment-for-OOD-generalization。
原文摘要 · Abstract (English)
Systematics contaminate observables, leading to distribution shifts relative to theoretically simulated signals-posing a major challenge for using pre-trained models to label such observables. Since systematics are often poorly understood and difficult to model, removing them directly and entirely may not be feasible. To address this challenge, we propose a novel method that aligns learned features between in-distribution (ID) and out-of-distribution (OOD) samples by optimizing a feature-alignment loss on the representations extracted from a pre-trained ID model. We first experimentally validate the method on the MNIST dataset using possible alignment losses, including mean squared error and optimal transport, and subsequently apply it to large-scale maps of neutral hydrogen. Our results show that optimal transport is particularly effective at aligning OOD features when parity between ID and OOD samples is unknown, even with limited data-mimicking real-world conditions in extracting information from large-scale surveys. Our code is available at https://github.com/sultan-hassan/feature-alignment-for-OOD-generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。