用SciBERT在512词长限制下高效分类天文论文,性能顶尖。
Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT
- 基于SciBERT的分类模型,处理512词上限文本。
- 宏F1达0.89,在WASP-2025竞赛中排名第一。
- 验证了领域预训练模型在短文本下的鲁棒性。
构建望远镜文献目录是评估观测台科学影响力和确保天文学可复现性的关键步骤,涉及识别、分类并关联引用或使用特定望远镜的科研论文。然而该过程仍高度依赖人工且资源消耗大。本文提出一种高效的SciBERT方法,实现四类自动分类:科学类、仪器类、提及类与非望远镜类。尽管面临最大512个标记(tokens)的严格上下文长度限制及有限算力,该方法仍取得0.89的宏F1得分,在WASP-2025排行榜中位居榜首。我们分析了截断的影响,发现即使一半样本超出长度限制,由于SciBERT的领域对齐特性,分类依然稳健。文章还探讨了截断、分块与长上下文模型之间的权衡,为科学文本整理提供了效率边界洞见。
原文摘要 · Abstract (English)
The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy. This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes. However, this process remains largely manual and resource intensive. In this work, we present an efficient SciBERT-based approach for automatic classification of scientific papers into four categories - science, instrumentation, mention, and not telescope. Despite strict context-length constraints (maximum 512 tokens) and limited compute resources, our approach achieved a macro F1 score of 0.89, ranking at the top of the WASP-2025 leaderboard. We analyze the effect of truncation and show that even with half the samples exceeding the token limit, SciBERT's domain alignment enables robust classification. We discuss trade-offs between truncation, chunking, and long-context models, providing insights into the efficiency frontier for scientific text curation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。