arXiv:2605.02035cs.CLcs.AI2026-05被引 1

构建视觉依赖性歧义数据集,提升多模态翻译的歧义消解能力

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

论文配图:VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
图 1 · 摘自论文原文
  • 设计2500个需视觉信息才能消歧的翻译样本
  • 提出基于LLM判别器的细粒度消歧评估指标
  • 验证思维链微调能更好处理未知歧义类型

歧义消解是多模态机器翻译(MMT)的关键挑战,模型必须真正利用视觉信息将模糊表达映射到其正确含义。尽管已有研究提出了面向歧义消解的评测基准,但现有基准仍受限于任务格式不匹配、歧义覆盖范围窄或视觉依赖性验证不足。此外,现有歧义评估方法难以适应开放文本中的多样歧义类型。为此,我们提出VIDA(Visually-Dependent Ambiguity),一个包含2,500个精心标注实例的数据集,其中每个标注源片段的消歧均需依赖视觉证据。我们进一步提出以消歧为核心的评估指标(Disambiguation-Centric Metrics),使用大语言模型作为裁判分类器,在词元层面验证歧义是否被正确消解。在两种先进视觉语言大模型上的实验表明,监督微调(SFT)提升了整体翻译质量,而思维链微调(CoT-SFT)在分布外歧义消解上表现更优,表明显式消歧引导有助于模型泛化至多样歧义类型。

原文摘要 · Abstract (English)

Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed disambiguation-oriented benchmarks probing the role of vision, we observe that existing benchmarks remain limited by task-format mismatch, narrow ambiguity coverage, or insufficient visual-dependency validation. Moreover, existing ambiguity evaluations are not well suited to diverse ambiguity types in open-ended translation. To address these limitations, we present VIDA (Visually-Dependent Ambiguity), a dataset of 2,500 carefully curated instances in which resolving an annotated source span requires visual evidence. We further propose Disambiguation-Centric Metrics that use an LLM-as-a-judge classifier to verify whether annotated ambiguous expressions are resolved correctly at the span level. Experiments with two state-of-the-art LVLMs show that supervised fine-tuning (SFT) improves overall translation quality, while chain-of-thought SFT (CoT-SFT) yields stronger out-of-distribution disambiguation, suggesting that explicit disambiguation guidance improves generalization to diverse ambiguity types.

多模态翻译歧义消解数据集视觉依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。