arXiv:2509.00664cs.CVcs.AI2025-09

用双编码器融合提升视觉感知,解决多模态模型看细节的短板

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

  • 用主编码器+增强编码器,通过轻量交叉注意力融合特征
  • 在TextVQA等5个基准上显著超越单编码器模型
  • 适合需要精细视觉理解的多模态应用开发

多模态大语言模型在连接视觉感知与高层文本推理方面取得显著进展,但存在根本矛盾:虽擅长复杂语义理解,却常在需精确细节感知的基础视觉任务中表现不佳。这主要源于现有架构普遍依赖单一视觉编码器,该编码器为高阶语义对齐优化,牺牲了细粒度视觉信息捕捉能力。为此,我们提出Fusion to Enhance(FtZ)新型视觉塔框架,突破单编码器设计,创新性地通过轻量级多头交叉注意力机制,将语义强大的锚点编码器与感知丰富的增强编码器组合。实验表明,在需细粒度视觉理解的多个挑战性基准(TextVQA、POPE、MMMU、MME 和 MM-Vet)上,我们的FtZ模型显著优于仅使用单编码器或现有特征融合方法的基线。本工作证明,组合异构专家编码器是克服当前多模态大模型视觉感知瓶颈的有效路径,为构建具备更强感知能力的下一代AI系统提供新范式。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made significant progress in bridging visual perception with high-level textual reasoning. However, they face a fundamental contradiction: while excelling at complex semantic understanding, these models often fail at basic visual tasks that require precise detail perception. This deficiency primarily stems from the prevalent architectural reliance on a single vision encoder optimized for high-level semantic alignment, which inherently sacrifices the ability to capture fine-grained visual information. To address this issue, we introduce Fusion to Enhance (FtZ), a novel vision tower framework. FtZ moves beyond the single-encoder design by innovatively composing a semantically powerful anchor encoder with a perception-rich augmenting encoder via a lightweight Multi-Head Cross-Attention mechanism. Experimental results demonstrate that on several challenging benchmarks demanding fine-grained visual understanding, such as TextVQA, POPE, MMMU, MME and MM-Vet, our FtZ model significantly outperforms baselines that use only a single encoder or existing feature fusion methods. This work proves that composing heterogeneous expert encoders is an efficient and effective path to overcoming the visual perception bottleneck in current MLLMs, offering a new design paradigm for building next-generation AI systems with stronger perceptual capabilities.

多模态视觉编码模型融合细粒度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。