无需图像也能精准翻译,通过图结构融合多模态信息
GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation
- 构建场景图保留视觉与语言特征,统一融合到共享空间
- 在无图像条件下仍达当前最优,英法/德翻译均超越基线
- 支持跨域推理,适合需要泛化能力的低资源翻译任务
多模态机器翻译(MMT)已证明视觉信息对翻译有显著帮助。然而现有方法受限于严格的视觉-语言对齐,在训练域内推理且难以处理模态差异。本文提出新型多模态场景图,以保留并整合模态特异性信息,并设计两阶段图引导归纳式无图像多模态翻译框架GIIFT。该框架采用跨模态图注意力网络适配器,在统一融合空间中学习多模态知识,并实现对更广范围无图像翻译任务的归纳推广。在Multi30K数据集上,英法、英语到德语任务均达到当前最优性能,且推理时无需图像。WMT基准测试结果表明,相比无图像翻译基线有显著提升,验证了GIIFT在归纳式无图像推理中的优势。
原文摘要 · Abstract (English)
Multimodal Machine Translation (MMT) has demonstrated the significant help of visual information in machine translation. However, existing MMT methods face challenges in leveraging the modality gap by enforcing rigid visual-linguistic alignment whilst being confined to inference within their trained multimodal domains. In this work, we construct novel multimodal scene graphs to preserve and integrate modality-specific information and introduce GIIFT, a two-stage Graph-guided Inductive Image-Free MMT framework that uses a cross-modal Graph Attention Network adapter to learn multimodal knowledge in a unified fused space and inductively generalize it to broader image-free translation domains. Experimental results on the Multi30K dataset of English-to-French and English-to-German tasks demonstrate that our GIIFT surpasses existing approaches and achieves the state-of-the-art, even without images during inference. Results on the WMT benchmark show significant improvements over the image-free translation baselines, demonstrating the strength of GIIFT towards inductive image-free inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。