让智能导航理解人的感受与反应,更像真人。
EgoCogNav: Cognition-aware Human Egocentric Navigation
- 融合视觉、视线和运动历史,预测人对路径的感知不确定性
- 在真实场景中实现轨迹与头部动作的精准预测
- 适合研究人类导航行为或开发智能助行系统的人
建模人类导航中的认知与体验因素,是深化人-环境交互理解、实现安全社交导航与有效辅助导引的核心。现有方法多聚焦于完全观测场景下的运动预测,常忽略影响人们感知空间的心理因素。为此,我们提出 EgoCogNav:一种多模态自指导航框架,可联合预测从第一视角视频、视线数据及运动历史中提取的路径感知不确定性、轨迹与头部运动。为推动该领域研究,我们构建了包含6小时真实世界第一视角记录的「认知感知自指导航(CEN)」数据集,涵盖多种真实场景下的导航行为。实验表明,EgoCogNav 学习到的感知不确定性与人类典型行为(如环视、犹豫、折返)高度相关,同时在未见导航记录上提升了轨迹与头部运动预测性能。
原文摘要 · Abstract (English)
Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding. Most existing methods focus on forecasting motions in fully observed scenes and often neglect human factors that capture how people feel and respond to space. To address this gap, we propose EgoCogNav, a multimodal egocentric navigation framework that jointly forecasts perceived path uncertainty, trajectories and head motion from egocentric video, gaze, and motion history. To facilitate research in the field, we introduce the Cognition-aware Egocentric Navigation (CEN) dataset consisting of 6 hours real-world egocentric recordings capturing diverse navigation behaviors in real-world scenarios. Experiments show that EgoCogNav learns the perceived uncertainty that strongly correlates with human-like behaviors such as scanning, hesitation, and backtracking while improving trajectory and head-motion forecasting on held-out navigation recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。