arXiv:2603.10613cs.CLcs.CV2026-03中稿 · LREC 2026被引 1

首个多语言新闻图像字幕数据集,覆盖9种语言。

MUNIChus: Multilingual News Image Captioning Benchmark

  • 构建跨语言新闻图文生成基准,含低资源语种
  • 现有模型在多语言场景下表现仍不理想
  • 适合研究多语言视觉-语言模型的学者使用

新闻图像字幕任务旨在结合新闻文本与对应图像生成描述,突出文本上下文与视觉元素的关系。目前大多数研究集中于英语,主要因其他语言数据集稀缺。为解决此问题,我们创建首个多语言新闻图像字幕基准MUNIChus,涵盖9种语言,包括僧伽罗语和乌尔都语等低资源语言。我们在MUNIChus上评估多种前沿神经新闻图像字幕模型,发现该任务在多语言环境下依然具有挑战性。MUNIChus已公开,包含超过20个模型的基准结果。该数据集为多语言新闻图像字幕模型的研发与评估开辟了新路径。

原文摘要 · Abstract (English)

The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and visual elements. The majority of research on news image captioning focuses on English, primarily because datasets in other languages are scarce. To address this limitation, we create the first multilingual news image captioning benchmark, MUNIChus, comprising 9 languages, including several low-resource languages such as Sinhala and Urdu. We evaluate various state-of-the-art neural news image captioning models on MUNIChus and find that news image captioning remains challenging. We also make MUNIChus publicly available with over 20 models already benchmarked. MUNIChus opens new avenues for further advancements in developing and evaluating multilingual news image captioning models.

多语言图像字幕新闻数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。