arXiv:2602.04600cs.RO2026-02

用人类第一视角数据训练机器人主动感知,提升复杂环境下的操作能力。

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data

  • 将主动感知建模为依赖历史的感知-动作循环,支持信息探索与决策分支。
  • 在轮式人形机器人上实现多场景长时序任务的强鲁棒性表现。
  • 融合人体运动与视觉时序信息,学习通用探索与操作先验,适合机器人具身智能研究者。

在开放环境中实现可泛化的操作能力,需要机器人主动解决信息不确定性,即具备主动感知能力。然而现有方法通常局限于有限的感知行为类型,限制了其在复杂环境中的适用性。本文将主动感知形式化为由信息探索动作和决策分支驱动的历史依赖型感知-动作循环,并提供视觉主动感知范式的结构化分类。基于此视角,提出 CoMe-VLA 框架——一种认知与记忆感知的视觉-语言-动作(VLA)模型,利用大规模人类第一人称数据学习多样化的探索与操作先验。该框架包含认知辅助头以实现自主子任务切换,以及双轨记忆系统,通过融合本体感受与视觉时序上下文,维持自我与环境的一致性感知。通过在统一的第一人称动作空间中对齐人类与机器人的手眼协调行为,采用三阶段渐进式训练。在轮式人形机器人上的大量实验表明,所提方法在跨越多种主动感知场景的多样化长时序任务中展现出强大的鲁棒性与适应性。

原文摘要 · Abstract (English)

Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confined in limited types of sensing behaviors, restricting their applicability to complex environments. In this work, we formalize active perception as a history-dependent perception-action loop driven by information-seeking action and decision branching, providing a structured categorization of visual active perception paradigms. Building on this perspective, we introduce CoMe-VLA, a cognitive and memory-aware vision-language-action (VLA) framework that leverages large-scale human egocentric data to learn versatile exploration and manipulation priors. Our framework integrates a cognitive auxiliary head for autonomous sub-task transitions and a dual-track memory system to maintain consistent self and environmental awareness by fusing proprioceptive and visual temporal contexts. By aligning human and robot hand-eye coordination behaviors in a unified egocentric action space, we train the model progressively in three stages. Extensive experiments on a wheel-based humanoid have demonstrated strong robustness and adaptability of our proposed method across diverse long-horizon tasks spanning multiple active perception scenarios.

主动感知具身智能视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。