arXiv:2501.08443cs.CVcs.LG2025-01被引 8

让视觉模型分层特征按指令动态融合,提升多模态任务表现

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models

  • 根据文本指令动态整合不同层级视觉特征
  • 在18个基准上验证,性能优于统一融合方法
  • 适合需要精细感知与语义理解的任务场景

大型视觉语言模型(LVLMs)通过整合预训练视觉编码器与大语言模型,在多种多模态任务中取得显著进展。然而,现有方法主要依赖视觉编码器最后一层的特征,忽略了浅层特征中蕴含的互补信息。尽管近期研究尝试使用多层视觉特征,但大多缺乏任务针对性,未考察层次化特征与具体任务间的依赖关系。为此,我们系统性地在18个涵盖6类任务的基准上,分析不同编码层特征的贡献。结果表明,多层特征具有互补优势,且其价值随任务变化,而统一融合会导致性能下降。基于此,我们提出指令引导的视觉聚合模块,可根据文本指令动态融合多层视觉特征,无需增加视觉标记数量。大量实验验证了该方法的优越性。进一步分析显示,中高阶特征在语义丰富任务中占主导,低阶特征对细粒度感知至关重要。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable success in a wide range of multimodal tasks by integrating pre-trained vision encoders and large language models. However, current LVLMs primarily rely on visual features extracted from the final layers of the vision encoder, overlooking the complementary information available in shallower layers. While recent approaches have explored the use of multilayer visual features in LVLMs, they tend to be task-agnostic and fail to examine the dependencies of hierarchical visual features on specific tasks. To address these gaps, we systematically investigate the contributions of visual features from different encoder layers using 18 benchmarks spanning 6 task categories. Our findings reveal that multilayer features provide complementary strengths with varying task dependencies, and uniform fusion leads to suboptimal performance. Building on these insights, we propose the instruction-guided vision aggregator, a module that dynamically integrates multi-layer visual features based on textual instructions, without increasing the number of visual tokens. Extensive evaluations demonstrate the superior performance of our method. Additionally, an in-depth analysis of the aggregator's behavior highlights the dominance of mid-to-high-level features in semantic-rich tasks and the critical role of low-level features in fine-grained perception.

多模态视觉语言特征融合指令引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。