用大模型自动给科技文献打标签,支持英德双语
SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
- 基于大模型构建多语言标签推荐系统
- 集成模型与合成数据提升标签准确率
- 适合数字图书馆智能化分类的实践参考
我们介绍 SemEval-2025 Task 5: LLMs4Subjects,一项关于使用 GND 分类体系对英文和德文科学与技术记录进行自动化主题标注的共享任务。参赛者开发基于大模型的系统,为每条记录推荐 top-k 主题标签,并通过定量指标(精确率、召回率、F1 值)和领域专家的定性评估进行评测。结果表明,大模型集成、合成数据生成以及多语言处理方法显著提升了标注效果,为大模型在数字图书馆分类中的应用提供了重要洞见。
原文摘要 · Abstract (English)
We present SemEval-2025 Task 5: LLMs4Subjects, a shared task on automated subject tagging for scientific and technical records in English and German using the GND taxonomy. Participants developed LLM-based systems to recommend top-k subjects, evaluated through quantitative metrics (precision, recall, F1-score) and qualitative assessments by subject specialists. Results highlight the effectiveness of LLM ensembles, synthetic data generation, and multilingual processing, offering insights into applying LLMs for digital library classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。