提出PRe方法缓解多模态大模型视觉表征退化问题
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
- 用预测初始视觉特征来正则中间层表示
- 显著提升模型在视觉语言任务上的表现
- 适合关注多模态模型内部视觉能力的研究者
尽管多模态大语言模型(MLLMs)在视觉-语言任务上表现优异,但其以语言为导向的训练对内部视觉基础能力的影响尚不明确。本文通过细致诊断发现,与初始视觉特征相比,大型语言模型中层的视觉表征存在全局功能和局部块结构的退化。我们归因于单一文本生成目标导致的视觉牺牲——模型为优化答案生成而降低视觉保真度。我们认为,一个鲁棒的MLLM需兼具跨模态推理与核心视觉能力,因此提出预测正则化(PRe),强制退化的中间特征预测初始视觉特征,从而保持内部表示的固有视觉属性。大量实验表明,缓解视觉退化能有效提升视觉-语言性能,凸显在MLLM中构建稳健内部视觉表征对全面多模态理解的关键作用。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradation in MLLMs. Specifically, we find that compared to the initial visual features, the visual representation in the middle layers of LLM exhibits both a degradation in global function and patch structure. We attribute this phenomenon to a visual sacrifice driven by the singular text-generation objective, where the model compromises its visual fidelity to optimize for answer generation. We argue that a robust MLLM requires both strong cross-modal reasoning and core visual competence, and propose Predictive Regularization (PRe) to force degraded intermediate features to predict initial visual features, thereby maintaining the inherent visual attributes of the MLLM's internal representations. Extensive experiments confirm that mitigating this visual degradation effectively boosts vision-language performance, underscoring the critical importance of fostering robust internal visual representations within MLLMs for comprehensive multimodal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。