用语义对齐提升图书馆数据的主题标注准确率
Homa at SemEval-2025 Task 5: Aligning Librarian Records with OntoAligner for Subject Tagging
- 将主题标注转化为记录与分类体系的语义对齐任务
- 在多语言环境下实现高精度标签匹配,效果优于传统方法
- 适合数字图书馆、知识管理领域研究者参考
本文介绍Homa系统在SemEval-2025任务5中的表现,该任务旨在基于Gemeinsame Normdatei(GND)分类体系,自动为TIBKAT技术记录分配主题标签。我们采用OntoAligner这一模块化本体对齐工具,结合检索增强生成(RAG)技术,将主题标注问题建模为语义相似性对齐任务。通过匹配记录与GND类别,评估了OntoAligner在主题索引中的适应性,并分析其处理多语言记录的能力。实验结果验证了该方法的有效性,也揭示了其局限性,表明对齐技术在数字图书馆主题标注中具有显著潜力。
原文摘要 · Abstract (English)
This paper presents our system, Homa, for SemEval-2025 Task 5: Subject Tagging, which focuses on automatically assigning subject labels to technical records from TIBKAT using the Gemeinsame Normdatei (GND) taxonomy. We leverage OntoAligner, a modular ontology alignment toolkit, to address this task by integrating retrieval-augmented generation (RAG) techniques. Our approach formulates the subject tagging problem as an alignment task, where records are matched to GND categories based on semantic similarity. We evaluate OntoAligner's adaptability for subject indexing and analyze its effectiveness in handling multilingual records. Experimental results demonstrate the strengths and limitations of this method, highlighting the potential of alignment techniques for improving subject tagging in digital libraries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。