arXiv:2410.02027cs.CVcs.AI2024-10EMNLP被引 6

对比德语原生描述与翻译描述,发现多语言视觉模型性能差距

Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

  • 用德语原生描述和翻译描述对比训练模型
  • 翻译版本导致平均召回率下降1.3%
  • 适合关注多语言跨文化差异的研究者

目前缺乏能充分考虑不同语言文化中图像描述感知差异的多语言视觉-语言模型。本文通过一个多模态、多语言检索案例研究,量化了现有模型灵活性的不足。实证表明,使用德语原生感知描述训练的模型,与使用英语翻译成德语的机器或人工翻译描述相比,存在显著性能差距。为此,我们进一步提出并评估了描述增强策略,虽实现了平均召回率提升1.3%,但差距依然存在,提示该领域仍是未来研究的重要方向。

原文摘要 · Abstract (English)

There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, multilingual retrieval case study, we quantify the existing lack of model flexibility. We empirically show performance gaps between training on captions that come from native German perception and captions that have been either machine-translated or human-translated from English into German. To address these gaps, we further propose and evaluate caption augmentation strategies. While we achieve mean recall improvements (+1.3), gaps still remain, indicating an open area of future work for the community.

多语言视觉语言检索跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。