用大模型自动分类天文望远镜文献,提升科研影响力评估效率
amc: The Automated Mission Classifier for Telescope Bibliographies
- 基于大语言模型自动识别论文中的望远镜引用
- 在TRACS挑战赛上达0.84宏F1得分,表现优异
- 可发现历史数据标签错误,适合天文档案与情报研究
望远镜文献库通过记录出版统计与引文指标,反映天文学研究的动态。稳健且可扩展的文献库有助于衡量设施与档案的科学影响。然而,出版物增长速度已超过人工标注能力。为此,我们提出自动化任务分类器(amc),利用大语言模型处理大量论文文本,自动识别并分类望远镜引用。amc改进版在TRACS Kaggle挑战赛中,于留出测试集上取得0.84的宏F1分数。该工具不仅适用于TRACS,还用于识别美国宇航局任务相关成果论文。此外,我们探索了amc在历史数据核查中的应用,可揭示潜在标签错误。结果表明,基于大语言模型的应用为图书馆科学提供了强大且可扩展的辅助手段。
原文摘要 · Abstract (English)
Telescope bibliographies record the pulse of astronomy research by capturing publication statistics and citation metrics for telescope facilities. Robust and scalable bibliographies ensure that we can measure the scientific impact of our facilities and archives. However, the growing rate of publications threatens to outpace our ability to manually label astronomical literature. We therefore present the Automated Mission Classifier (amc), a tool that uses large language models (LLMs) to identify and categorize telescope references by processing large quantities of paper text. A modified version of amc performs well on the TRACS Kaggle challenge, achieving a macro $F_1$ score of 0.84 on the held-out test set. amc is valuable for other telescopes beyond TRACS; we developed the initial software for identifying papers that featured scientific results by NASA missions. Additionally, we investigate how amc can also be used to interrogate historical datasets and surface potential label errors. Our work demonstrates that LLM-based applications offer powerful and scalable assistance for library sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。