从天文文献中挖掘历史关联,预测未来新概念-对象关系。
Predicting New Concept-Object Associations in Astronomy by Mining the Literature
- 构建天文文献知识图谱,自动提取天体与科学概念
- 基于历史结构的矩阵分解模型在预测上提升16.8% (NDCG@100)
- 适合需要高效筛选观测目标的研究者使用
我们通过自动化流程从截至2025年7月的astro-ph全文语料中构建了一个概念-对象知识图谱。该流程对光学字符识别(OCR)处理后的论文进行命名实体抽取,将天体识别为SIMBAD标识符,并将其链接至原始语料中标注的科学概念。随后测试历史图结构是否能提前预测新出现的概念-对象关联。由于概念来自聚类,存在语义重叠,因此所有方法均在推理阶段统一采用概念相似性平滑。在四个时间截点下,针对物理意义明确的概念子集,带有平滑的隐式反馈矩阵分解模型(交替最小二乘法,ALS)在NDCG@100上比最强的邻域基线(使用文本嵌入相似性的KNN)高出16.8%(0.144 vs 0.123),Recall@100提高19.8%(0.175 vs 0.146),分别超过最佳时效性启发式方法96%和88%。结果表明,历史文献中蕴含了全局启发式或局部邻域投票未捕捉的预测结构,为稀缺望远镜时间的后续目标筛选提供了新路径。
原文摘要 · Abstract (English)
We construct a concept-object knowledge graph from the full astro-ph corpus through July 2025. Using an automated pipeline, we extract named astrophysical objects from OCR-processed papers, resolve them to SIMBAD identifiers, and link them to scientific concepts annotated in the source corpus. We then test whether historical graph structure can forecast new concept-object associations before they appear in print. Because the concepts are derived from clustering and therefore overlap semantically, we apply an inference-time concept-similarity smoothing step uniformly to all methods. Across four temporal cutoffs on a physically meaningful subset of concepts, an implicit-feedback matrix factorization model (alternating least squares, ALS) with smoothing outperforms the strongest neighborhood baseline (KNN using text-embedding concept similarity) by 16.8% on NDCG@100 (0.144 vs 0.123) and 19.8% on Recall@100 (0.175 vs 0.146), and exceeds the best recency heuristic by 96% and 88%, respectively. These results indicate that historical literature encodes predictive structure not captured by global heuristics or local neighborhood voting, suggesting a path toward tools that could help triage follow-up targets for scarce telescope time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。