arXiv:2602.09934cs.CV2026-02被引 3

让大模型视觉主干同时胜任语义理解与像素级任务

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

  • 用多任务协同微调优化视觉主干,引入细粒度监督信号
  • 在语义分割、深度估计等密集预测任务上显著提升性能
  • 适合需要兼顾语言推理与像素级分析的多模态应用

多模态大语言模型(MLLM)在视觉-语言理解中取得显著进展,其视觉编码器展现出优越的高层语义对齐能力。一个关键问题随之而来:这些编码器能否作为通用视觉主干,可靠完成经典视觉任务?为此,我们做出以下贡献:(i) 发现MLLM中的视觉编码器在密集特征表示上存在缺陷,表现为在密集预测任务(如语义分割、深度估计)上表现不佳;(ii) 提出VersaViT,一种基于新型多任务框架的视觉变压器,通过轻量级任务头与多粒度监督实现协同后训练;(iii) 在多种下游任务上的广泛实验验证了方法的有效性,成功构建了一个兼具语言引导推理与像素级理解能力的通用视觉主干。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within their vision encoders. An important question thus arises: Can these encoders serve as versatile vision backbones, capable of reliably performing classic vision-centric tasks as well? To address the question, we make the following contributions: (i) we identify that the vision encoders within MLLMs exhibit deficiencies in their dense feature representations, as evidenced by their suboptimal performance on dense prediction tasks (e.g., semantic segmentation, depth estimation); (ii) we propose VersaViT, a well-rounded vision transformer that instantiates a novel multi-task framework for collaborative post-training. This framework facilitates the optimization of the vision backbone via lightweight task heads with multi-granularity supervision; (iii) extensive experiments across various downstream tasks demonstrate the effectiveness of our method, yielding a versatile vision backbone suited for both language-mediated reasoning and pixel-level understanding.

视觉主干多任务学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。