用对比学习将X射线谱与文献知识对齐,提升天体物理数据解读效率。
Augmenting representations with scientific papers
- 设计对比学习框架,对齐X射线谱与科学文献中的领域知识。
- 谱图检索文献的Recall@1%达20%,且20个物理变量估计提升16-18%。
- 适用于罕见源发现,可推广至其他需结合文献的科学领域。
天文学积累了海量多模态数据(图像、光谱、时序),并有数十年文献分析天体源。但这些数据与文献极少系统整合。本文提出一种对比学习框架,将X射线光谱与从科学文献中提取的领域知识对齐,构建共享多模态表征。由于文献涵盖更广的物理背景,对齐极具挑战。该框架在谱图检索文献任务中达到20% Recall@1%,证明跨模态对齐可行且能加速稀有或难理解源的解读。共享隐空间有效编码物理意义信息,融合光谱与文本数据后,20个物理变量估计精度较单模态光谱基线提升16-18%。混合专家(MoE)策略利用单模态与共享表征,表现更优。多模态隐空间中的异常分析识别出高优先级目标,包括候选脉动超亮X射线源(PULX)和引力透镜系统。该框架可拓展至其他需对齐观测数据与文献的科学领域。
原文摘要 · Abstract (English)
Astronomers have acquired vast repositories of multimodal data, including images, spectra, and time series, complemented by decades of literature that analyzes astrophysical sources. Still, these data sources are rarely systematically integrated. This work introduces a contrastive learning framework designed to align X-ray spectra with domain knowledge extracted from scientific literature, facilitating the development of shared multimodal representations. Establishing this connection is inherently complex, as scientific texts encompass a broader and more diverse physical context than spectra. We propose a contrastive pipeline that achieves a 20% Recall@1% when retrieving texts from spectra, proving that a meaningful alignment between these modalities is not only possible but capable of accelerating the interpretation of rare or poorly understood sources. Furthermore, the resulting shared latent space effectively encodes physically significant information. By fusing spectral and textual data, we improve the estimation of 20 physical variables by 16-18% over unimodal spectral baselines. Our results indicate that a Mixture of Experts (MoE) strategy, which leverages both unimodal and shared representations, yields superior performance. Finally, outlier analysis within the multimodal latent space identifies high-priority targets for follow-up investigation, including a candidate pulsating ULX (PULX) and a gravitational lens system. Importantly, this framework can be extended to other scientific domains where aligning observational data with existing literature is possible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。