arXiv:2605.12960cs.CL2026-05

不重新训练,让多模态模型轻松学会57种语言

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

论文配图:DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
图 1 · 摘自论文原文
  • 通过方向与幅度感知融合,精准整合多语言与多模态更新
  • 在57种语言上显著提升多语言性能,保持原有多模态能力
  • 适合想低成本扩展多语言能力的模型开发者

为实现更通用的人类级智能,大语言模型需无缝融合多语言与多模态能力;然而,将现有多模态模型扩展至多种语言通常需要昂贵的多语言多模态数据构建和重复端到端微调。本文研究一种无需训练的替代方案:通过在共享语言模型主干中组合残差更新,向现有多模态模型注入多语言能力。核心挑战在于多语言与多模态更新异质,反映共享模型中不同功能角色。为此,提出方向与幅度感知的多语言多模态融合(DiM3),在每个参数维度选择性地融合两类更新,同时保留原始视觉编码器和多模态投影器。在基于LLaVA和Qwen的多语言基准测试中,涵盖57种语言的文本与视觉-语言场景下,DiM3持续优于现有融合基线,显著提升多语言性能,并在保持通用多模态能力的同时,达到与专用多语言多模态微调相当的效果。进一步实验表明,DiM3可直接应用于已训练的多语言多模态模型,仍能带来额外增益。可解释性分析显示,DiM3主要重塑中间层语义表征,在文本与多模态输入下强化跨语言对齐,同时保留高层任务敏感结构。

原文摘要 · Abstract (English)

Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; however, extending an existing multimodal model to many languages typically requires expensive multilingual multimodal data construction and repeated end-to-end retraining. We study a training-free alternative: injecting multilingual capability into an existing multimodal model by composing residual updates in the shared language model backbone. The key challenge is that multilingual and multimodal updates are heterogeneous, reflecting different functional roles in the shared model. To address this, we propose Direction- and Magnitude-aware Multilingual Multimodal merging (DiM3), which selectively composes the two updates at each parameter dimension while preserving the original vision encoder and multimodal projector. Experiments on multilingual benchmarks in both text-only and vision-language settings, covering 57 languages across LLaVA- and Qwen-based backbones, show that DiM3 consistently outperforms existing merging baselines, substantially improves multilingual performance over the original multimodal model, and remains competitive with dedicated multilingual multimodal fine-tuning while largely retaining general multimodal ability. We further show that DiM3 can be directly applied to already trained multilingual multimodal models and still yield additional gains. Further interpretability analysis shows that DiM3 primarily reshapes intermediate-layer semantic representations, strengthening cross-lingual alignment under both text-only and multimodal inputs while preserving higher-layer task-sensitive structure. Our repository is on https://github.com/wzj1718/DiM3.

多语言多模态模型融合迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。