arXiv:2506.17516cs.ROcs.CV2025-06中稿 · ICRA被引 3

无需标注和外部奖励,实现动态事件感知与主动追踪

EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization

  • 通过最小化自由能量,统一时空表征与具身控制
  • 利用预测误差和熵作为内在信号,自动分割事件并追踪目标
  • 适合无人工标注的动态场景,如机器人协作与自主导航

主动事件感知能力对于人机协作、辅助机器人和自主导航等任务中的具身智能至关重要。然而,现有方法通常依赖预定义动作空间、标注数据集和外在奖励,限制了其在动态真实场景中的适应性和可扩展性。受事件感知认知理论和预测编码启发,我们提出 EASE——一种基于自监督能量最小化的框架,通过自由能量最小化统一时空表征学习与具身控制。EASE 利用预测误差和熵作为内在信号,实现事件分割、观测摘要与主动追踪,无需显式标注或外部奖励。通过将生成感知模型与动作驱动控制策略耦合,EASE 动态对齐预测与观测,涌现出隐式记忆、目标连续性及对新环境的适应能力。仿真与真实世界评估表明,EASE 能实现隐私保护且可扩展的事件感知,为非脚本化动态任务中的具身系统提供稳健基础。

原文摘要 · Abstract (English)

Active event perception, the ability to dynamically detect, track, and summarize events in real time, is essential for embodied intelligence in tasks such as human-AI collaboration, assistive robotics, and autonomous navigation. However, existing approaches often depend on predefined action spaces, annotated datasets, and extrinsic rewards, limiting their adaptability and scalability in dynamic, real-world scenarios. Inspired by cognitive theories of event perception and predictive coding, we propose EASE, a self-supervised framework that unifies spatiotemporal representation learning and embodied control through free energy minimization. EASE leverages prediction errors and entropy as intrinsic signals to segment events, summarize observations, and actively track salient actors, operating without explicit annotations or external rewards. By coupling a generative perception model with an action-driven control policy, EASE dynamically aligns predictions with observations, enabling emergent behaviors such as implicit memory, target continuity, and adaptability to novel environments. Extensive evaluations in simulation and real-world settings demonstrate EASE's ability to achieve privacy-preserving and scalable event perception, providing a robust foundation for embodied systems in unscripted, dynamic tasks.

具身智能事件感知自监督学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。