用大词汇量标签集做德国学术文献自动主题标引,生成式AI表现更优。
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature

- 对比了监督式XMLC与基于LLM的生成方法在主题标引中的效果。
- 生成式方法在长尾标签和专业评分中胜出,整体准确率超85%。
- 适合需要高覆盖度的主题标引场景,如图书馆知识管理。
在拥有大量受控词汇作为标签集的情况下,图书馆中的自动化主题标引可视为多标签分类任务。若主题词数量庞大,则符合极端多标签分类(XMLC)的目标。本研究将一系列专用监督式XMLC方法应用于德国国家图书馆(DNB)收集的当代德语科学文献主题标引任务。同时引入经典词汇匹配基线及三种近期开发的基于大语言模型(LLM)的方法进行基准测试。所有算法在多个指标下评估,包括与已有索引材料的二值相关性比较,以及专业主题馆员的分级相关性评分。所有方法均面临从主题词汇长尾中可靠提出建议的挑战。结果表明,依赖于基于Transformer的密集特征的监督式XMLC算法在整体二值相关性指标上表现最佳;但若聚焦分级相关性及长尾标签表现,基于LLM的生成方法表现更优,显示出未来实际应用的潜力。
原文摘要 · Abstract (English)
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。