arXiv:2510.14307cs.CLcs.AI2025-10Transactions of th…

构建多语言多模态实体链接测试平台,提升跨语言图文实体识别准确率。

MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

  • 构建五种语言图文数据集,覆盖7000+实体提及与2500+维基数据实体
  • 视觉信息可显著提升模糊文本中的实体链接准确率,尤其对弱多语模型有效
  • 开源数据集与基准方法,适合多模态、跨语言研究者使用

本文提出MERLIN,一个用于多语言多模态实体链接任务的新测试平台。该数据集包含来自BBC新闻文章标题及其对应图像的五种语言(印地语、日语、印尼语、越南语、泰米尔语)数据,涵盖超过7,000个命名实体提及,关联至2,500个唯一的Wikidata实体。我们还引入多个基于多语言和多模态实体链接方法的基准测试,使用LLaMa-2和Aya-23等语言模型进行评估。结果表明,融入视觉信息能显著提高实体链接精度,尤其是在文本上下文模糊或不足的情况下,且对多语能力较弱的模型效果尤为明显。相关数据集、代码与方法已公开于https://github.com/rsathya4802/merlin。

原文摘要 · Abstract (English)

This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hindi, Japanese, Indonesian, Vietnamese, and Tamil, featuring over 7,000 named entity mentions linked to 2,500 unique Wikidata entities. We also include several benchmarks using multilingual and multimodal entity linking methods exploring different language models like LLaMa-2 and Aya-23. Our findings indicate that incorporating visual data improves the accuracy of entity linking, especially for entities where the textual context is ambiguous or insufficient, and particularly for models that do not have strong multilingual abilities. For the work, the dataset, methods are available here at https://github.com/rsathya4802/merlin

多模态实体链接多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。