用视觉场景图剪枝提升多模态翻译准确性
Multimodal Machine Translation with Visual Scene Graph Pruning
- 基于语言信息剪裁视觉场景图冗余节点
- 在多个数据集上显著提升翻译准确率
- 适合关注视觉信息融合的NLP研究者
多模态机器翻译(MMT)通过引入视觉信息应对语言多义性和歧义问题。当前研究的主要瓶颈在于如何有效利用视觉数据。以往方法通常提取图像全局或区域特征,通过注意力或门控机制融合多模态信息,但未能充分解决视觉信息冗余问题。本文提出一种新方法——视觉场景图剪枝(PSG),利用语言场景图信息指导视觉场景图中冗余节点的剔除,从而降低下游翻译任务中的噪声。通过与前沿方法的大量对比实验及消融研究,验证了PSG模型的有效性。结果表明,视觉信息剪枝在推动MMT发展方面具有巨大潜力。
原文摘要 · Abstract (English)
Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bottleneck in current MMT research is the effective utilization of visual data. Previous approaches have focused on extracting global or region-level image features and using attention or gating mechanisms for multimodal information fusion. However, these methods have not adequately tackled the issue of visual information redundancy in MMT, nor have they proposed effective solutions. In this paper, we introduce a novel approach--multimodal machine translation with visual Scene Graph Pruning (PSG), which leverages language scene graph information to guide the pruning of redundant nodes in visual scene graphs, thereby reducing noise in downstream translation tasks. Through extensive comparative experiments with state-of-the-art methods and ablation studies, we demonstrate the effectiveness of the PSG model. Our results also highlight the promising potential of visual information pruning in advancing the field of MMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。