arXiv:2506.11515cs.CVcs.CL2025-06中稿 · IEEE Transactions …被引 1

提出轻量级模块Manager,提升双塔模型跨模态对齐效果

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

  • 引入Manager模块,动态聚合多层级单模态专家知识
  • 在4个下游任务上超越现有基线,零样本性能显著提升
  • 适用于双塔VLM与MLLM,适配多种图像分辨率和能力类型

双塔视觉-语言模型(VLM)在多个下游任务中表现优异。尽管BridgeTower通过构建编码器间桥梁提升了性能,但仍存在三方面问题:(i)未能有效利用单模态表示的逐层信息,(ii)限制了不同层次单模态语义知识的灵活使用,(iii)仅限于传统低分辨率数据集上的评估。本文提出Manager,一个轻量、高效且有效的插件,可自适应聚合预训练单模态专家在不同层级提供的洞察,促进更全面的跨模态对齐与融合。在双塔VLM架构下,我们设计ManagerTower,在每个跨模态层引入管理器,无论是否进行视觉语言预训练,其在4个下游任务上均优于先前强基线。进一步扩展至最新多模态大语言模型(MLLM)架构,结果表明LLaVA-OV-Manager显著提升LLaVA-OV在20个下游数据集上的零样本性能,涵盖不同能力类别、图像与分辨率,无论是否启用多网格算法。深入分析显示,经理模块与多网格算法均可视为插件,通过捕捉深度与宽度两个正交视角的多样化视觉细节来增强视觉表征。二者协同可缓解多网格算法带来的语义模糊,进一步提升性能。代码与模型已开源。

原文摘要 · Abstract (English)

Two-Tower Vision--Language Models (VLMs) have demonstrated strong performance across various downstream VL tasks. While BridgeTower further enhances performance by building bridges between encoders, it \textit{(i)} suffers from ineffective layer-by-layer utilization of unimodal representations, \textit{(ii)} restricts the flexible exploitation of different levels of unimodal semantic knowledge, and \textit{(iii)} is limited to the evaluation on traditional low-resolution datasets only with the Two-Tower VLM architecture. In this work, we propose Manager, a lightweight, efficient and effective plugin that adaptively aggregates insights from different levels of pre-trained unimodal experts to facilitate more comprehensive VL alignment and fusion. First, under the Two-Tower VLM architecture, we introduce ManagerTower, a novel VLM that introduces the manager in each cross-modal layer. Whether with or without VL pre-training, ManagerTower outperforms previous strong baselines and achieves superior performance on 4 downstream VL tasks. Moreover, we extend our exploration to the latest Multimodal Large Language Model (MLLM) architecture. We demonstrate that LLaVA-OV-Manager significantly boosts the zero-shot performance of LLaVA-OV across different categories of capabilities, images, and resolutions on 20 downstream datasets, whether the multi-grid algorithm is enabled or not. In-depth analysis reveals that both our manager and the multi-grid algorithm can be viewed as a plugin that improves the visual representation by capturing more diverse visual details from two orthogonal perspectives (depth and width). Their synergy can mitigate the semantic ambiguity caused by the multi-grid algorithm and further improve performance. Code and models are available at https://github.com/LooperXX/ManagerTower.

双塔模型跨模态对齐MLLM插件设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。