针对天文光变曲线不规则采样难题,提出多视角自蒸馏学习方法。
Domain-Informed Multi-View Self-Distillation for Astronomical Light-Curve Representation Learning with JEPA

- 融合语义保持视图与不确定感知分词,用多视角自蒸馏提升表征能力。
- 在星体分类任务中,1样本/类时宏F1达42.56,100样本/类时达63.58。
- 适用于变星识别、参数估计等场景,适合天文学与不规则时间序列研究者。
光变曲线描述天体亮度随时间的变化。学习鲁棒的光变曲线表征对大规模动态宇宙自动发现至关重要,但现有时间序列基础模型常面临不均匀采样、复杂噪声及宽泛物理时标等挑战。本文提出一种面向不规则天文时间序列的领域感知表示学习框架,基于联合嵌入预测架构(JEPA),结合语义保持视图、不确定感知分词与多视角自蒸馏。编码器在LEAVES数据集上通过LeJEPA正则化训练,并在StarEmbed分类基准上评估。在StarEmbed上,该模型在16项分类指标中有15项优于手工特征;少样本线性探针下,每类仅1样本时宏F1为42.56±7.21,每类100样本时达63.58±1.20,持续优于手工特征。除变星分类外,所学表征支持相似性搜索、参数估计与光度零点漂移检测。跨域适应在12个来自PYRREGULAR的异构不规则时间序列数据集上评估,适配版本在5个数据集上匹配或超越此前最优表现,而任一先前基线最多仅胜3个。结果表明,领域感知的多视角自蒸馏是学习不规则时间序列表征的有效策略,同时强调成功的时间序列表征学习需依赖领域特定归纳偏置而非通用最优架构。
原文摘要 · Abstract (English)
Light curves describe temporal variations in the brightness of celestial objects. Learning robust representations of light curves is essential for large-scale automatic discovery in the dynamic universe, but existing time-series foundation models often struggle with the uneven sampling, complex noise, and wide range of physical timescales that characterize astronomical observations. We propose a domain-informed representation learning framework for irregular astronomical time series with Joint-Embedding predictive architecture (JEPA), combining semantics-preserving views, uncertainty-aware tokenization, and multi-view self-distillation. The encoders are trained with multi-view self-distillation using LeJEPA regularization on the LEAVES dataset and evaluated on the StarEmbed classification benchmark. On StarEmbed, our model outperforms hand-crafted features on 15 of 16 classification metrics. In few-shot linear probing, it achieves macro-F1 scores of 42.56 $\pm$ 7.21 with one sample per class and 63.58 $\pm$ 1.20 with 100 samples per class, consistently improving over hand-crafted features. Beyond variable-star classification, the learned representation supports similarity search, parameter estimation, and photometric zero-point drift detection. We further evaluate cross-domain adaptation on 12 heterogeneous irregular time-series datasets from PYRREGULAR, where the adapted variant matches or exceeds previous state-of-the-art performance on 5 datasets, compared with at most 3 wins by any single prior baseline. These results demonstrate that domain-informed multi-view self-distillation is an effective strategy for learning representations of irregular time series, while also highlighting that successful time-series representation learning requires domain-specific inductive biases rather than a universally optimal architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。