arXiv:2503.02505cs.AIcs.CV2025-03被引 12

用户用自己视角的分割图指挥智能体,让其在3D环境里更准地完成任务。

ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment

  • 用跨视角对齐框架,让用户通过自身摄像头的分割图指定目标。
  • 推理效率比前代提升3到6倍,且能在不同环境中零样本泛化。
  • 适合希望简化人机交互、提升智能体空间理解能力的研究者。

我们旨在开发一种语义清晰、空间敏感、领域无关且对用户直观的目标指定方法,以引导智能体在3D环境中的交互。为此,提出一种新颖的跨视角目标对齐框架,允许用户使用自身相机视图中的分割掩码指定目标对象,而非依赖智能体的观测。我们指出,当人类与智能体的视角差异显著时,仅靠行为克隆无法对齐智能体行为与人类意图。为此,引入两种辅助目标:跨视角一致性损失和目标可见性损失,以显式增强智能体的空间推理能力。基于此,我们构建了ROCKET-2,一个在Minecraft中训练的先进智能体,在推理效率上相比ROCKET-1提升3至6倍。实验表明,ROCKET-2可直接解析来自人类视角的目标,显著改善人机交互。值得注意的是,尽管仅在Minecraft数据集上训练,它仍能通过简单的动作空间映射,实现对Doom、DMLab和Unreal等其他3D环境的零样本泛化。

原文摘要 · Abstract (English)

We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent's observations. We highlight that behavior cloning alone fails to align the agent's behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop ROCKET-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3x to 6x compared to ROCKET-1. We show that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping.

人机交互视觉导航零样本泛化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。