arXiv:2504.21589cs.CLcs.AI2025-04ACL被引 10

用大模型集成实现文献自动主题标引,兼顾精度与专家评价。

DNB-AI-Project at SemEval-2025 Task 5: An LLM-Ensemble Approach for Automated Subject Indexing

  • 通过少样本提示多个大模型生成关键词,结合后处理映射与投票。
  • 在全部主题赛道中量化排名第四,专家质评排名第一。
  • 适合需要高精度主题标引的图书馆与知识管理场景。

本文介绍了为 SemEval-2025 任务 5:LLMs4Subjects(基于大模型的国家技术图书馆开放馆藏自动主题标签)所开发的系统。该系统通过向多个不同能力的大语言模型输入已标注记录的少量示例,要求其对新记录进行关键词建议。采用少样本提示方法,并结合一系列后处理步骤:将生成关键词映射至目标词汇表、聚合结果进行集成投票,最后按相关性排序。系统在全主题赛道的量化排名中位列第四,但在由主题标引专家进行的定性评估中取得最佳成绩。

原文摘要 · Abstract (English)

This paper presents our system developed for the SemEval-2025 Task 5: LLMs4Subjects: LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog. Our system relies on prompting a selection of LLMs with varying examples of intellectually annotated records and asking the LLMs to similarly suggest keywords for new records. This few-shot prompting technique is combined with a series of post-processing steps that map the generated keywords to the target vocabulary, aggregate the resulting subject terms to an ensemble vote and, finally, rank them as to their relevance to the record. Our system is fourth in the quantitative ranking in the all-subjects track, but achieves the best result in the qualitative ranking conducted by subject indexing experts.

主题标引大模型集成少样本学习信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。