让视觉信息从配角变主角,提升多模态理解精度
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
- 将视觉信号作为预测目标而非输入条件,统一建模视觉与语言
- 在通用任务和视觉主导任务上均表现优异,无需额外设计
- 为构建全能型视觉智能体提供新范式,适合多模态研究者
尽管视觉语言模型(VLMs)取得显著进展,但现有架构常因难以保留细粒度视觉信息而造成粗粒度的多模态理解。我们归因于当前训练范式存在文本主导的优化偏差,将视觉信号视为被动的条件输入而非监督目标。为此,我们提出 Youtu-VL 框架,采用视觉-语言统一自回归监督(VLUAS)范式,从根本上将优化目标从“视觉作为输入”转变为“视觉作为目标”。通过将视觉标记直接融入预测序列,Youtu-VL 对视觉细节与语言内容施加统一的自回归监督。此外,该范式可扩展至以视觉为中心的任务,使标准 VLM 无需任务特定改造即可完成视觉主导任务。大量实验证明,Youtu-VL 在通用多模态任务与视觉主导任务上均达到竞争力水平,为通用视觉智能体的发展奠定坚实基础。
原文摘要 · Abstract (English)
Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, leading to coarse-grained multimodal comprehension. We attribute this deficiency to a suboptimal training paradigm inherent in prevailing VLMs, which exhibits a text-dominant optimization bias by conceptualizing visual signals merely as passive conditional inputs rather than supervisory targets. To mitigate this, we introduce Youtu-VL, a framework leveraging the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm, which fundamentally shifts the optimization objective from ``vision-as-input'' to ``vision-as-target.'' By integrating visual tokens directly into the prediction stream, Youtu-VL applies unified autoregressive supervision to both visual details and linguistic content. Furthermore, we extend this paradigm to encompass vision-centric tasks, enabling a standard VLM to perform vision-centric tasks without task-specific additions. Extensive empirical evaluations demonstrate that Youtu-VL achieves competitive performance on both general multimodal tasks and vision-centric tasks, establishing a robust foundation for the development of comprehensive generalist visual agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。