首个多语言图书主题词标注基准,支持精确与概念匹配评估。
LCSHBench: A Multilingual, Consensus-Grounded Benchmark for Library of Congress Subject Heading Assignment
- 基于三所高校22,346本书的共识标注数据,确保主题一致性。
- 跨语言检索在exact recall@200达0.659,优于3072维模型的0.623。
- 适合研究多语言信息检索、主题标注与语义对齐的学者使用。
自动化主题编目为书目记录分配受控词汇主题词,但美国国会图书馆主题词表(LCSH)缺乏公开标准基准。我们提出LCSHBench:涵盖哈佛、哥伦比亚和普林斯顿三所高校开放许可馆藏的22,346本书,涉及15种语言。仅当至少两个独立编目机构赋予相同主题词时才纳入记录,并提供各馆来源信息及并集与一致答案视图。对465,187项三馆共同编目的汇编研究显示:图书馆通常在概念层级上达成一致(93.3%共享概念级主题词),但在具体表达上差异明显(39.4%主题词集合完全相同)。因此,LCSHBench同时评估精确匹配与概念匹配,按语言和主题类型分解集合与排序指标,覆盖开放词汇生成与全词汇检索任务。初步实验表明,一个300M参数的轻量级嵌入模型经低秩微调后,跨语言检索表现更优,在开发集exact recall@200达到0.659,超过3,072维云端嵌入模型的0.623。语言层面分析显示提升非均匀分布,留出测试与端到端验证待后续工作。
原文摘要 · Abstract (English)
Automated subject cataloging assigns controlledvocabulary headings to bibliographic records, but LCSH has no standard public benchmark. We introduce LCSHBench: 22,346 books in 15 languages from the openly licensed Harvard, Columbia, and Princeton catalogs. Records enter only when at least two independent cataloging agencies assigned LCSH; we release per-catalog provenance plus union and unanimous answer views. A concordance study of 465,187 works cataloged by all three libraries shows why this design matters: libraries usually agree on the underlying topic (93.3% share a concept-level heading) but often differ in exact expression (39.4% have identical heading sets). LCSHBench therefore scores both exact and concept matches, with set and rank metrics broken down by language and heading type, across open-vocabulary generation and full-vocabulary retrieval. As a first demonstration, a low-rank fine-tune of a 300M on-device embedder improves cross-lingual retrieval and beats a 3,072-dimensional hosted embedder on development exact recall@200 (0.659 vs 0.623). The language panel shows the gain is not uniform, and held-out-test and end-to-end confirmation remain future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。