arXiv:2507.02997cs.LGcs.AI2025-07被引 1

从第一视角视频中学习规划动作,通过记忆环境特征提升执行鲁棒性。

What to Do Next? Memorizing skills from Egocentric Instructional Video

  • 用拓扑可达性记忆结合Transformer建模环境结构
  • 在动作偏离时仍能保持高成功率与稳定表现
  • 适合需要环境理解的机器人自主决策场景

通过示范学习行为需从观察中提取有意义的环境信息。本研究在仿真环境中从第一人称视角探讨高层目标导向动作的规划问题。提出一项新任务——交互式动作规划,并设计一种结合拓扑可达性记忆与Transformer架构的方法。通过提取环境中的可操作特征来记忆环境结构,从而根据上下文选择合适动作。此外,该记忆模型可在完成特定目标时检测动作偏差。为验证方法的通用性,我们在一个真实感交互仿真环境中进行评估。实验结果表明,所提方法能学习到有意义的表征,在动作出现偏差时仍表现出更高性能和更强鲁棒性。

原文摘要 · Abstract (English)

Learning to perform activities through demonstration requires extracting meaningful information about the environment from observations. In this research, we investigate the challenge of planning high-level goal-oriented actions in a simulation setting from an egocentric perspective. We present a novel task, interactive action planning, and propose an approach that combines topological affordance memory with transformer architecture. The process of memorizing the environment's structure through extracting affordances facilitates selecting appropriate actions based on the context. Moreover, the memory model allows us to detect action deviations while accomplishing specific objectives. To assess the method's versatility, we evaluate it in a realistic interactive simulation environment. Our experimental results demonstrate that the proposed approach learns meaningful representations, resulting in improved performance and robust when action deviations occur.

动作规划第一视角记忆机制仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。