构建天体物理多标签分类数据集,解决术语极端不平衡问题。
AstroConcepts: A Large-Scale Multi-Label Classification Corpus for Astrophysics
- 基于21,702篇论文抽象,标注2,367个统一天文学术语
- 76%概念训练样本少于50,揭示极端类别不平衡现象
- 提出频率分层评估法,暴露聚合指标掩盖的性能差异
科学多标签文本分类面临严重类别不平衡,专业术语呈现极端幂律分布,挑战常规分类方法。现有科学语料库缺乏全面受控词汇表,多聚焦宽泛类别,限制对极端不平衡的系统研究。本文引入AstroConcepts,包含21,702篇已发表天体物理论文的英文摘要,使用统一天文学术语表(Unified Astronomy Thesaurus)标注2,367个概念。该语料库表现出严重标签不平衡,76%的概念训练样本少于50条。通过发布此资源,我们支持科学领域极端不平衡的系统研究,并在传统方法、神经网络及词汇约束大模型上建立强基线。评估揭示三个关键模式:第一,词汇约束大模型在天体物理分类中表现接近领域适配模型,提示参数高效方法潜力;第二,领域适配对罕见专业术语提升显著,但所有方法绝对性能仍有限;第三,提出频率分层评估,揭示聚合评分掩盖的性能规律,使鲁棒性评估成为科学多标签分类的核心。这些结果为科学NLP提供可操作洞见,并建立极端不平衡研究基准。
原文摘要 · Abstract (English)
Scientific multi-label text classification suffers from extreme class imbalance, where specialized terminology exhibits severe power-law distributions that challenge standard classification approaches. Existing scientific corpora lack comprehensive controlled vocabularies, focusing instead on broad categories and limiting systematic study of extreme imbalance. We introduce AstroConcepts, a corpus of English abstracts from 21,702 published astrophysics papers, labeled with 2,367 concepts from the Unified Astronomy Thesaurus. The corpus exhibits severe label imbalance, with 76% of concepts having fewer than 50 training examples. By releasing this resource, we enable systematic study of extreme class imbalance in scientific domains and establish strong baselines across traditional, neural, and vocabulary-constrained LLM methods. Our evaluation reveals three key patterns that provide new insights into scientific text classification. First, vocabulary-constrained LLMs achieve competitive performance relative to domain-adapted models in astrophysics classification, suggesting a potential for parameter-efficient approaches. Second, domain adaptation yields relatively larger improvements for rare, specialized terminology, although absolute performance remains limited across all methods. Third, we propose frequency-stratified evaluation to reveal performance patterns that are hidden by aggregate scores, thereby making robustness assessment central to scientific multi-label evaluation. These results offer actionable insights for scientific NLP and establish benchmarks for research on extreme imbalance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。