通过精准微调特定语言层,高效提升视觉语言模型多语言能力。
Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models
- 识别浅层中与语言相关的神经元激活,定位多语言理解关键层。
- 仅微调14%参数,多语言性能显著提升,跨语言任务准确率更高。
- 适用于低资源和复杂视觉推理,适合追求效率的多语言模型优化。
大型视觉语言模型(LVLMs)在理解视觉信息方面表现出色,但在多语言能力上存在不平衡。本文深入研究了LVLMs的多语言工作模式,发现其多语言理解能力与浅层中的语言特异性神经元激活密切相关。基于此,提出PLAST训练方案:通过监测语言特异性神经元激活,识别多语言理解相关层,并利用问题-翻译对进行精准微调,实现多语言对齐。在MM-Bench和MMMB上的实验表明,PLAST能有效提升LVLM的多语言能力,且仅需微调14%参数即达成显著效率。进一步分析显示,该方法可推广至低资源及复杂视觉推理任务,促进浅层中语言特异性视觉信息的交互。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have demonstrated exceptional capabilities in understanding visual information with human languages but also exhibit an imbalance in multilingual capabilities. In this work, we delve into the multilingual working pattern of LVLMs and identify a salient correlation between the multilingual understanding ability of LVLMs and language-specific neuron activations in shallow layers. Building on this insight, we introduce PLAST, a training recipe that achieves efficient multilingual enhancement for LVLMs by Precise LAnguage-Specific layers fine-Tuning. PLAST first identifies layers involved in multilingual understanding by monitoring language-specific neuron activations. These layers are then precisely fine-tuned with question-translation pairs to achieve multilingual alignment. Our empirical results on MM-Bench and MMMB demonstrate that PLAST effectively improves the multilingual capabilities of LVLMs and achieves significant efficiency with only 14% of the parameters tuned. Further analysis reveals that PLAST can be generalized to low-resource and complex visual reasoning tasks, facilitating the language-specific visual information engagement in shallow layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。