arXiv:2510.24285cs.CVcs.AI2025-10被引 9

让视觉语言模型自我进化,提升细粒度视觉感知能力。

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model

  • 构建粗到细的渐进式学习任务,实现自批判与自预测的闭环训练。
  • 在7个基准上平均提升1.7%,细粒度任务最高提升6.0%。
  • 适用于追求视觉感知能力持续进化的AI研发人员。

视觉语言模型(VLMs)在真实应用中受限于细粒度视觉感知能力不足。现有方法面临高质量数据稀缺和训练策略局限:监督微调(SFT)常损害通用能力,强化学习微调(RFT)则偏重文本推理。为此,我们提出一种两阶段任务,将视觉感知学习建模为从粗到细的渐进过程。基于此,我们开发了ViPER——一个自启动框架,通过图像级与实例级重建结合两阶段强化学习策略,实现自生成数据驱动的闭环训练。应用于Qwen2.5-VL系列,生成Qwen-Viper系列模型,在涵盖多种任务的七个综合基准上平均提升1.7%,细粒度感知任务最高提升6.0%,性能全面优于基线并保持通用性。该框架不仅实现感知能力的自我提升,还验证了生成与理解间的相互促进关系,为构建更自主、更强的VLMs提供新范式。

原文摘要 · Abstract (English)

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often compromises general capabilities, while reinforcement fine-tuning (RFT) prioritizes textual reasoning over visual perception. To bridge this gap, we propose a novel two-stage task that structures visual perception learning as a coarse-to-fine progressive process. Based on this task formulation, we develop ViPER, a self-bootstrapping framework specifically designed to enable iterative evolution through self-critiquing and self-prediction. By synergistically integrating image-level and instance-level reconstruction with a two-stage reinforcement learning strategy, ViPER establishes a closed-loop training paradigm, where internally synthesized data directly fuel the enhancement of perceptual ability. Applied to the Qwen2.5-VL family, ViPER produces the Qwen-Viper series. With an average gain of 1.7% on seven comprehensive benchmarks spanning various tasks and up to 6.0% on fine-grained perception, Qwen-Viper consistently demonstrates superior performance across different vision-language scenarios while maintaining generalizability. Beyond enabling self-improvement in perceptual capabilities, ViPER provides concrete evidence for the reciprocal relationship between generation and understanding, a breakthrough to developing more autonomous and capable VLMs.

视觉语言模型自进化细粒度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。