用扩散模型为句子生成想象图,提升翻译准确率
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
- 用扩散模型自动生成与句意一致的图像
- 在Multi30K上翻译准确率提升超14个BLEU点
- 适合想用视觉信息增强翻译的研究者
视觉信息已被用于提升机器翻译性能,但其效果高度依赖大量带人工标注图像的双语平行语料。本文将基于稳定扩散的想象网络引入多模态大语言模型(MLLM),为每个源句显式生成对应图像,推动多模态机器翻译发展。特别地,我们通过强化学习构建启发式人类反馈机制,确保生成图像与源句语义一致,无需图像标注监督,突破了视觉信息在翻译中应用的瓶颈。此外,该方法还能将想象的视觉信息融入大规模纯文本翻译任务。实验表明,所提模型显著优于现有多模态及纯文本翻译方法,尤其在Multi30K多模态翻译基准上平均提升超过14个BLEU点。
原文摘要 · Abstract (English)
Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we introduce a stable diffusion-based imagination network into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing the multimodel MT. Particularly, we build heuristic human feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of image annotation, which breaks the bottleneck of using visual information in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into large-scale text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 14 BLEU points on Multi30K multimodal MT benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。