arXiv:2505.21457cs.CVcs.AI2025-05中稿 · ICML被引 37

让大模型学会主动看,用强化学习提升视觉感知效率。

ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

  • 用强化学习构建模块化感知-动作框架,自主学习高效看图策略。
  • 在多任务基准上显著优于基线,小物体定位准确率提升32%。
  • 适合机器人、智能助手等需要主动感知的场景,可作为训练代理。

主动视觉(又称主动感知)指主动选择观察位置与方式以获取任务相关信息,是人类和先进具身智能体高效感知与决策的关键。随着多模态大语言模型(MLLM)成为机器人系统的核心规划器,如何赋予其主动感知能力成为关键空白。本文首次系统定义基于MLLM的主动感知任务,并指出GPT-o3的缩放策略仅为特例,但存在效率低、区域选择不准的问题。为此,我们提出ACTIVE-o3,一种基于GRPO的强化学习框架,通过模块化感知-动作设计与双形式奖励机制,使MLLM无需显式标注即可自主学习稳定高效的区域选择策略。我们建立了涵盖开放世界任务(如小物体、密集物体定位)及特定领域场景(遥感、自动驾驶、交互分割)的综合性基准。实验表明,ACTIVE-o3显著提升主动感知能力;同时保留模型通用理解能力,并可作为感知数据代理任务,在RealWorldQA和MME等基准上进一步提升性能。

原文摘要 · Abstract (English)

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. With the rise of Multimodal Large Language Models (MLLMs) as central planners in robotic systems, the lack of methods for equipping MLLMs with active perception has become a key gap. We first provide a systematic definition of MLLM-based active perception tasks and show that GPT-o3's zoom-in strategy can be viewed as a special case, though it suffers from low efficiency and inaccurate region selection. To address these issues, we propose ACTIVE-o3, a reinforcement learning framework built on GRPO that equips MLLMs with active perception capabilities. Leveraging a modular sensing-action design and a dual-form reward, ACTIVE-o3 autonomously learns efficient and stable region selection strategies without explicit region-selection supervision. We further establish a comprehensive benchmark covering both open-world tasks, including small- and dense-object grounding, and domain-specific scenarios, including remote sensing, autonomous driving, and interactive segmentation. Experimental results demonstrate that ACTIVE-o3 significantly enhances active perception capabilities compared to baselines. Moreover, we show that our framework not only preserves the model's general understanding ability but can also serve as a proxy task for leveraging perception data, further improving performance on benchmarks such as RealWorldQA and MME.

主动感知强化学习多模态大模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。