构建多语言书目索引数据集,支持权威词表驱动的智能编目。
An Extreme Multi-label Text Classification (XMTC) Library Dataset: What if we took "Use of Practical AI in Digital Libraries" seriously?
- 构建英德双语书目数据集,含权威词表(GND)标注
- 实现基于权威词表的多标签分类,准确率超基线模型
- 适合图书馆学与AI交叉研究者,推动可解释的智能编目
主题索引对信息发现至关重要,但大规模跨语言维护困难。本文发布一个大型英德双语书目记录语料库,包含整合权威文件(GND)的标注,并提供可机器读取的GND分类体系。该资源支持基于本体的多标签分类,实现文本到权威术语的映射,以及可复现、基于权威的辅助编目评估。我们提供了三个系统的统计概览和定性错误分析。诚邀社区不仅评估准确率,更关注实用性与透明度,共同推进以权威为锚点的AI协作者,助力编目员工作增效。
原文摘要 · Abstract (English)
Subject indexing is vital for discovery but hard to sustain at scale and across languages. We release a large bilingual (English/German) corpus of catalog records annotated with the Integrated Authority File (GND), plus a machine-actionable GND taxonomy. The resource enables ontology-aware multi-label classification, mapping text to authority terms, and agent-assisted cataloging with reproducible, authority-grounded evaluation. We provide a brief statistical profile and qualitative error analyses of three systems. We invite the community to assess not only accuracy but usefulness and transparency, toward authority-anchored AI co-pilots that amplify catalogers' work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。