arXiv:2507.20411cs.CL2025-07被引 2

用图像概念增强检索,让多语言图文生成更准更省数据

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning

  • 用图像概念+检索片段联合生成caption,避免依赖英译句
  • 在低资源语言上表现接近高资源语言,数据需求大幅降低
  • 适合做多语言视觉理解、少样本场景下的模型开发

多语言视觉-语言模型在图像描述任务上已取得显著进展,但仍因多语言训练数据有限和大规模参数化成本高而落后于英文模型。检索增强生成(RAG)通过在目标语言中检索示例来生成描述,减少对多语言训练的依赖。然而现有方法常依赖从英语翻译的检索描述,易引入语义偏差和语言偏见。本文提出CONCAP,将检索到的描述与图像特定概念结合,提升输入图像的上下文感知能力,实现跨语言的精准描述生成。在XM3600数据集上的实验表明,CONCAP在低资源和中等资源语言上均表现优异,且显著降低数据需求。结果验证了概念感知的检索增强在缩小多语言性能差距方面的有效性。

原文摘要 · Abstract (English)

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameterization. Retrieval-augmented generation (RAG) offers a promising alternative by conditioning caption generation on retrieved examples in the target language, reducing the need for extensive multilingual training. However, multilingual RAG captioning models often depend on retrieved captions translated from English, which can introduce mismatches and linguistic biases relative to the source language. We introduce CONCAP, a multilingual image captioning model that integrates retrieved captions with image-specific concepts, enhancing the contextualization of the input image and grounding the captioning process across different languages. Experiments on the XM3600 dataset indicate that CONCAP enables strong performance on low- and mid-resource languages, with highly reduced data requirements. Our findings highlight the effectiveness of concept-aware retrieval augmentation in bridging multilingual performance gaps.

多语言生成检索增强图像描述概念融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。