arXiv:2505.12081cs.CV2025-05被引 14

一个模型统一解决视觉感知任务,靠强化学习生成可靠推理过程。

VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning

  • 用统一奖励机制和多目标认知学习提升模型推理能力。
  • 在检测、分割、计数任务上分别领先基线29.1%、22.1%、13.2%。
  • 无需标注推理数据,人类评估显示其推理过程可信可靠。

大视觉语言模型具备处理多种视觉感知任务的内在能力。本文提出VisionReasoner,一种统一框架,可在共享模型中完成多重视觉感知任务的推理与求解。通过设计统一奖励机制和多目标认知学习策略,VisionReasoner增强推理能力,对视觉输入进行分析,并在统一模型中应对多样感知任务。该模型在输出前生成结构化推理流程。人类评估表明,即使无标注推理训练数据,VisionReasoner的推理过程依然忠实可靠。为严格评估统一视觉感知能力,我们在涵盖检测、分割、计数三大领域的十项任务上评估VisionReasoner。实验结果表明,作为统一模型,其表现显著优于基线Qwen2.5VL:在COCO(检测)上提升29.1%,在ReasonSeg(分割)上提升22.1%,在CountBench(计数)上提升13.2%。

原文摘要 · Abstract (English)

Large vision-language models exhibit inherent capabilities to handle diverse visual perception tasks. In this paper, we introduce VisionReasoner, a unified framework capable of reasoning and solving multiple visual perception tasks within a shared model. Specifically, by designing a unified reward mechanism and multi-object cognitive learning strategies, VisionReasoner enhances its reasoning capabilities to analyze visual inputs, and addresses diverse perception tasks within a unified model. VisionReasoner generates a structured reasoning process before delivering the desired outputs responding to user queries. Human evaluation reveals the reasoning process of VisionReasoner is faithful and reliable even without annotated reasoning train data. To rigorously assess unified visual perception capabilities, we evaluate VisionReasoner on ten diverse tasks spanning three critical domains: detection, segmentation, and counting. Experimental results show that VisionReasoner achieves superior performance as a unified model, outperforming the baseline Qwen2.5VL by relative margins of 29.1\% on COCO (detection), 22.1\% on ReasonSeg (segmentation), and 13.2\% on CountBench (counting).

视觉推理统一模型强化学习多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。