arXiv:2505.23179cs.CV2025-05被引 11

用强化学习提升大模型对复杂场景的细粒度感知能力

DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes

  • 设计三类基于规则的奖励机制,引导模型分步审视复杂视觉场景
  • 在密集人群等挑战性场景中显著提升目标检测准确率
  • 适合关注视觉推理与不确定性建模的研究者和开发者

多模态大模型虽具备较强视觉理解能力,但在密集人群等复杂真实场景中的细粒度感知仍受限。受大语言模型中强化学习成功的启发,本文提出基于强化学习的深度检视与感知框架DIP-R1,旨在增强多模态大模型对复杂场景的感知能力。DIP-R1通过三种规则化奖励机制引导模型进行细致观察:首先采用标准推理奖励,推动模型完成‘整体理解→聚焦模糊区域→决策判断’三步推理;其次设计方差引导注视奖励,鼓励模型关注不确定区域,实现区域级不确定性推理并可解释地标识模糊区域;第三构建加权精确率-召回率奖励,提升决策准确性。在包含密集人群等挑战性真实场景的多个细粒度检测数据集上验证,DIP-R1在多种域内与域外场景中均实现一致且显著优于现有基线及SFT方法的性能。结果表明,将强化学习融入多模态大模型,在复杂现实感知任务中具有巨大潜力。

原文摘要 · Abstract (English)

MLLMs have demonstrated significant visual understanding capabilities, yet their fine-grained visual perception in complex real-world scenarios, such as densely crowded public areas, remains limited. Inspired by the recent success of RL in both LLMs and MLLMs, in this paper, we explore how RL can enhance visual perception ability of MLLMs. Then we develop a novel RL-based framework, Deep Inspection and Perception with RL (DIP-R1) designed to enhance the visual perception capabilities of MLLMs, by comprehending complex scenes and looking through visual instances closely. DIP-R1 guides MLLMs through detailed inspection of visual scene via three simply designed rule-based reward modeling. First, we adopt a standard reasoning reward encouraging the model to include three-step reasoning process: 1) comprehending entire visual scene, 2) observing for looking through interested but ambiguous regions, and 3) decision-making for predicting answer. Second, a variance-guided looking reward is designed to encourage MLLM to examine uncertain regions during the observing process, guiding it to inspect ambiguous areas and mitigate perceptual uncertainty. This reward promotes variance-driven visual exploration, enabling MLLM to reason about region-level uncertainty and explicitly indicate interpretable uncertain regions. Third, we model a weighted precision-recall accuracy reward enhancing accurate decision-making. We verify its effectiveness across diverse fine-grained object detection data consisting of challenging real-world scenes, such as densely crowded scenes. Built upon existing MLLMs, DIP-R1 achieves consistent and significant improvement across various in-domain and out-of-domain scenarios, outperforming various existing baselines and SFT method. Our findings highlight the substantial potential of integrating RL into MLLMs for enhancing capabilities in complex real-world perception tasks.

多模态强化学习视觉感知细粒度检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。