arXiv:2504.19675cs.CLcs.AI2025-04ACL被引 11

将传统主题词标引与大模型结合,提升多语言文献主题分类效果。

Annif at SemEval-2025 Task 5: Traditional XMTC augmented by LLMs

  • 融合传统NLP与大模型翻译、合成数据生成技术
  • 在全主题类别中排名第一,核心主题类第二
  • 适合需要高效多语言主题标引的图书馆与知识系统

本文介绍Annif系统在SemEval-2025任务5(LLMs4Subjects)中的表现,该任务聚焦于利用大语言模型(LLMs)进行主题标引。任务要求基于双语TIBKAT数据库的书目记录,使用GND主题词表生成主题预测。我们的方法结合了Annif工具包中传统的自然语言处理与机器学习技术,以及创新的基于大模型的翻译与合成数据生成方法,并整合单语模型的预测结果。系统在定量评估中,全主题类别排名第一,tib-core-subjects类别排名第二;在定性评估中位列第四。结果表明,将传统跨语言主题词标引(XMTC)算法与现代大模型技术相结合,可显著提升多语言环境下主题标引的准确性和效率。

原文摘要 · Abstract (English)

This paper presents the Annif system in SemEval-2025 Task 5 (LLMs4Subjects), which focussed on subject indexing using large language models (LLMs). The task required creating subject predictions for bibliographic records from the bilingual TIBKAT database using the GND subject vocabulary. Our approach combines traditional natural language processing and machine learning techniques implemented in the Annif toolkit with innovative LLM-based methods for translation and synthetic data generation, and merging predictions from monolingual models. The system ranked first in the all-subjects category and second in the tib-core-subjects category in the quantitative evaluation, and fourth in qualitative evaluations. These findings demonstrate the potential of combining traditional XMTC algorithms with modern LLM techniques to improve the accuracy and efficiency of subject indexing in multilingual contexts.

主题标引大模型多语言信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。