对比预训练编码器与解码器在多模态翻译中的效果,发现解码器预训练更有效。
Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation
- 系统测试不同预训练策略对多模态翻译的影响
- 预训练解码器显著提升翻译流畅性与准确率
- 适合关注多模态模型设计的研究者参考
多模态机器翻译(MMT)通过结合图像等辅助模态来提升翻译质量。尽管大规模预训练语言与视觉模型在单模态任务中表现优异,但其在多模态翻译中的作用仍不明确。本文在统一的MMT框架下,系统研究了预训练编码器与解码器的影响。实验基于Multi30K和CoMMuTE数据集,涵盖英德、英法翻译任务。结果表明:预训练在多模态场景中作用显著且不对称——预训练解码器始终带来更流畅、更准确的输出;而预训练编码器的效果则依赖于视觉-文本对齐质量。此外,本文揭示了模态融合与预训练组件之间的相互作用,为未来多模态翻译系统架构设计提供指导。
原文摘要 · Abstract (English)
Multimodal Machine Translation (MMT) aims to improve translation quality by leveraging auxiliary modalities such as images alongside textual input. While recent advances in large-scale pre-trained language and vision models have significantly benefited unimodal natural language processing tasks, their effectiveness and role in MMT remain underexplored. In this work, we conduct a systematic study on the impact of pre-trained encoders and decoders in multimodal translation models. Specifically, we analyze how different training strategies, from training from scratch to using pre-trained and partially frozen components, affect translation performance under a unified MMT framework. Experiments are carried out on the Multi30K and CoMMuTE dataset across English-German and English-French translation tasks. Our results reveal that pre-training plays a crucial yet asymmetrical role in multimodal settings: pre-trained decoders consistently yield more fluent and accurate outputs, while pre-trained encoders show varied effects depending on the quality of visual-text alignment. Furthermore, we provide insights into the interplay between modality fusion and pre-trained components, offering guidance for future architecture design in multimodal translation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。