arXiv:2505.24456cs.CL2025-05EMNLP被引 5

用图像提升机器翻译的文化敏感度,实测有效。

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

  • 构建5800组图文对,评估视觉信息对翻译的辅助作用
  • 加入图像后,文化专有项、性别标记等翻译准确率提升
  • 适合研究跨文化翻译与多模态模型的学者使用

由于文化间概念化差异,仅靠语言难以传递地域性语义,导致机器翻译在处理文化内容时面临挑战。本文探究图像是否可作为多模态翻译中的文化上下文。我们提出CaMMT,一个包含超过5800组图像与英-区域语言平行描述的高质量人工标注数据集。基于该数据集,我们评估了五种视觉语言模型(VLMs)在纯文本与图文结合两种设置下的表现。通过自动评估与人工评测发现,视觉上下文能普遍提升翻译质量,尤其在处理文化专有项(CSIs)、歧义消解和正确性别标记方面效果显著。通过公开发布CaMMT,旨在推动更贴近文化细节与地区差异的多模态翻译系统研发与评估。

原文摘要 · Abstract (English)

Translating cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meanings. In this work, we investigate whether images can act as cultural context in multimodal translation. We introduce CaMMT, a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. Using this dataset, we evaluate five Vision Language Models (VLMs) in text-only and text+image settings. Through automatic and human evaluations, we find that visual context generally improves translation quality, especially in handling Culturally-Specific Items (CSIs), disambiguation, and correct gender marking. By releasing CaMMT, our objective is to support broader efforts to build and evaluate multimodal translation systems that are better aligned with cultural nuance and regional variations.

多模态翻译文化敏感视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。