arXiv:2509.23957cs.CLcs.AI2025-09被引 1

让机器翻译借助视觉信息提升准确性

Vision-Grounded Machine Interpreting: Improving the Translation Process through Visual Cues

  • 融合语音与摄像头画面,用视觉上下文辅助翻译
  • 视觉信息显著提升词义消歧,对性别判断有小幅度提升
  • 适合需要场景理解的实时翻译任务

当前机器口译系统为单模态实时语音转语音架构,仅依赖语言信号进行翻译。这种单一模态依赖在需借助视觉、情境或语用信息进行消歧和语义准确性的场景中受限。本文提出视觉引导口译(Vision-Grounded Interpreting, VGI),构建一个集成视觉-语言模型的原型系统,通过网络摄像头同时处理语音与视觉输入,利用上下文视觉信息引导翻译过程。为评估该方法有效性,我们手工构建了一个针对三类歧义的诊断语料库。实验结果表明:视觉引导显著改善词义消歧;对性别指代解析带来适度但不稳定的提升;对句法歧义则无明显增益。研究认为,引入多模态是提升机器口译翻译质量的必要方向。

原文摘要 · Abstract (English)

Machine Interpreting systems are currently implemented as unimodal, real-time speech-to-speech architectures, processing translation exclusively on the basis of the linguistic signal. Such reliance on a single modality, however, constrains performance in contexts where disambiguation and adequacy depend on additional cues, such as visual, situational, or pragmatic information. This paper introduces Vision-Grounded Interpreting (VGI), a novel approach designed to address the limitations of unimodal machine interpreting. We present a prototype system that integrates a vision-language model to process both speech and visual input from a webcam, with the aim of priming the translation process through contextual visual information. To evaluate the effectiveness of this approach, we constructed a hand-crafted diagnostic corpus targeting three types of ambiguity. In our evaluation, visual grounding substantially improves lexical disambiguation, yields modest and less stable gains for gender resolution, and shows no benefit for syntactic ambiguities. We argue that embracing multimodality represents a necessary step forward for advancing translation quality in machine interpreting.

机器口译多模态视觉引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。