arXiv:2604.16943cs.CL2026-04中稿 · SCIS被引 2

通过识别神经元角色,精准微调多模态大模型以提升图像翻译准确率。

MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation

论文配图:MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
图 1 · 摘自论文原文
  • 基于指令激活分析定位视觉与语言模块中的关键神经元。
  • 仅微调语言相关和通用神经元,性能超越现有方法。
  • 适合需要高精度图像翻译的多模态应用开发者。

多模态大语言模型(MLLMs)虽具强大能力,却常难以捕捉图像中细微的文本信息,导致视觉输入与文本输出间的模态差距。现有方法依赖指令微调,易引发参数冗余,影响泛化性能。为此,本文提出模态神经元感知微调(MNAFT),通过指令驱动的激活分析,识别视觉与语言模块中语言无关与语言相关的神经元,并评估其在不同翻译任务中的重要性。随后,在目标任务相关层中仅对语言相关及通用神经元进行选择性微调,保留其他神经元与层的知识。大量实验表明,MNAFT显著优于当前最优图像翻译方法,包括级联模型、全量微调及参数高效微调技术。我们还通过神经元激活可视化与聚类分析,揭示不同神经元组在跨模态理解与语言特定翻译中的作用。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This often leads to a modality gap between visual text inputs and textual inputs/outputs for image translation. Existing methods, primarily relying on instruction fine-tuning, risk parameter redundancy of pre-trained knowledge, hindering generalization performance. To address this, we introduce modality neuron-aware fine-tuning (MNAFT), a novel approach that takes advantage of the specialized roles of individual neurons within MLLMs for enhanced image translation. MNAFT identifies language-agnostic and language-specific neurons in both vision and language modules through an instruction-driven activation analysis, evaluating their importance in various translation tasks. We then perform selective fine-tuning, updating only the parameters of language-specific and language-agnostic neurons within the selected layers relevant to the target task, while preserving the knowledge encoded in other neurons and layers. Our extensive experiments on multiple benchmarks demonstrate that MNAFT significantly outperforms state-of-the-art image translation methods, including cascaded models, standard full fine-tuning, and parameter-efficient tuning techniques. Furthermore, we provide comprehensive analysis, including visualizations of neuron activations and clustering patterns, to offer insights into the roles of different neuron groups in mediating cross-modal understanding and facilitating accurate language-specific translation.

多模态图像翻译微调神经元分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。