arXiv:2508.15877cs.CLcs.AI2025-08中稿 · KONVENS 2025被引 3

用高效小模型提升文献主题分类,获德国评测冠军

Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs

  • 采用小型高效模型生成翻译数据和候选主题
  • 结合大模型对主题候选进行排序,提升准确率
  • 适合追求效率与精度的学术主题自动标注场景

本文介绍了 Annif 系统在 GermEval-2025 LLMs4Subjects 共享任务(子任务2)中的表现。该任务要求利用大语言模型为书目记录生成主题标签,特别关注计算效率。我们的系统基于 Annif 自动主题索引工具包,改进了此前在首场 LLMs4Subjects 任务中表现优异的版本。通过使用多个小型高效语言模型完成翻译与合成数据生成,并利用大模型对候选主题进行排序,显著提升了性能。系统在定量评估和定性评估中均排名第一。

原文摘要 · Abstract (English)

This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using large language models, with a special focus on computational efficiency. Our system, based on the Annif automated subject indexing toolkit, refines our previous system from the first LLMs4Subjects shared task, which produced excellent results. We further improved the system by using many small and efficient language models for translation and synthetic data generation and by using LLMs for ranking candidate subjects. Our system ranked 1st in the overall quantitative evaluation of and 1st in the qualitative evaluation of Subtask 2.

主题分类高效模型文献标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。