通过选择性调制层与神经元,实现高效多语言多模态翻译
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
- 按语言特性选择关键层和神经元,减少冗余
- 仅微调40%参数即超越全量微调,性能达最优
- 适合资源有限下做多语言多模态翻译的场景
多模态机器翻译(MMT)通过引入视觉上下文提升翻译质量,有助于解决文本歧义。现有方法在双语场景表现良好,但扩展至多语言时面临跨语言干扰和参数共享效率低的问题。为此,我们提出LLaVA-NeuMT框架,显式建模语言特异与通用表示以缓解多语言干扰。方法包含层选择机制,识别不同语言对最有效的层;以及神经元级自适应策略,动态选择语言特异与通用神经元,提升翻译质量并降低冗余。在M3-Multi30K和M3-AmbigCaps数据集上实验表明,仅微调40%模型参数,其性能即超越全量微调方法,并在两个数据集上达到最先进水平。分析揭示了所选层与神经元在多模态多语言适配中的关键作用,为跨语言适配提供高效可扩展的解决方案。
原文摘要 · Abstract (English)
Multimodal Machine Translation (MMT) enhances translation quality by incorporating visual context, helping to resolve textual ambiguities. While existing MMT methods perform well in bilingual settings, extending them to multilingual translation remains challenging due to cross-lingual interference and ineffective parameter-sharing strategies. To address this, we propose LLaVA-NeuMT, a novel multimodal multilingual translation framework that explicitly models language-specific and language-agnostic representations to mitigate multilingual interference. Our approach consists of a layer selection mechanism that identifies the most informative layers for different language pairs and a neuron-level adaptation strategy that dynamically selects language-specific and agnostic neurons to improve translation quality while reducing redundancy. We conduct extensive experiments on the M3-Multi30K and M3-AmbigCaps datasets, demonstrating that LLaVA-NeuMT, while fine-tuning only 40\% of the model parameters, surpasses full fine-tuning approaches and ultimately achieves SOTA results on both datasets. Our analysis further provides insights into the importance of selected layers and neurons in multimodal multilingual adaptation, offering an efficient and scalable solution to cross-lingual adaptation in multimodal translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。