arXiv:2511.15279cs.ROcs.CV2025-11

让机器人像人一样主动看东西,用语言指令控制摄像头抓取最需要的信息。

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception

  • 统一建模视觉、语言和摄像头动作,实现端到端的主动感知。
  • 仅用500个真实样本训练,平均任务完成率达96%。
  • 适合需要精准视觉引导的机器人应用,如家务助手、巡检系统。

在具身智能中,视觉感知应是主动的:系统需自主决定看向何处、以何种尺度观测,以在像素和空间预算限制下获取最具信息量的数据。现有耦合固定RGB-D相机的视觉模型无法兼顾大范围覆盖与细粒度细节获取,严重制约其在开放世界机器人任务中的表现。本文研究语言引导的主动视觉感知任务:给定一张RGB图像和自然语言指令,智能体需输出真实PTZ(云台-俯仰-变焦)相机的平移、俯仰和缩放调整,以获取最相关信息。我们提出EyeVLA,一种统一框架,将视觉感知、语言理解与物理相机控制整合于单一自回归视觉-语言-动作模型中。EyeVLA引入语义丰富且高效的层次化动作编码,紧凑地对连续相机调整进行分词,并嵌入视觉语言模型词汇表,实现多模态联合推理。通过数据高效管道——包含伪标签生成、迭代交并比控制的数据精炼以及基于组相对策略优化(GRPO)的强化学习——我们仅用500个真实世界样本,便将预训练视觉语言模型的开放世界理解能力迁移到具身主动感知策略中。在50个不同真实场景、五次独立评估运行中,EyeVLA平均任务完成率达到96%。本工作建立了指令驱动的多模态具身系统主动视觉信息获取新范式。

原文摘要 · Abstract (English)

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models coupled with fixed RGB-D cameras fundamentally fail to reconcile wide-area coverage with fine-grained detail acquisition, severely limiting their efficacy in open-world robotic applications. We study the task of language-guided active visual perception: given a single RGB image and a natural language instruction, the agent must output pan, tilt, and zoom adjustments of a real PTZ (pan-tilt-zoom) camera to acquire the most informative view for the specified task. We propose EyeVLA, a unified framework that addresses this task by integrating visual perception, language understanding, and physical camera control within a single autoregressive vision-language-action model. EyeVLA introduces a semantically rich and efficient hierarchical action encoding that compactly tokenizes continuous camera adjustments and embeds them into the VLM vocabulary for joint multimodal reasoning. Through a data-efficient pipeline comprising pseudo-label generation, iterative IoU-controlled data refinement, and reinforcement learning with Group Relative Policy Optimization (GRPO), we transfer the open-world understanding of a pre-trained VLM to an embodied active perception policy using only 500 real-world samples. Evaluations on 50 diverse real-world scenes across five independent evaluation runs demonstrate that EyeVLA achieves an average task completion rate of 96%. Our work establishes a new paradigm for instruction-driven active visual information acquisition in multimodal embodied systems.

具身智能主动感知视觉语言模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。