用轻量模型预测人类注视轨迹,让机器人导航更像人
Fast Human Attention Prediction for Fixation-guided Active Perception in Autonomous Navigation

- 用液态神经网络+MobileNetV3实现低耗实时注视预测
- 仅0.61 GFLOPs却达最新基准,计算量降99.4%
- 已用于真实无人机导航,让机器视觉更贴近人类
人类视觉注意力依赖有结构的注视路径高效处理场景,但将此行为融入机器人自主系统仍处初期,受限于现有预测模型的高计算开销。为此,我们提出GazeLNN,一种计算轻量的注视路径预测模型,采用液态神经网络作为循环核心,结合MobileNetV3进行特征提取。该模型以自回归方式,基于当前视觉刺激和注视历史预测序列化注视热图。尽管仅需0.61 GFLOPs,GazeLNN在MIT低分辨率数据集上达到0.47的ScanMatch分数,优于现有循环基线模型,且计算成本降低99.40%,推理速度提升至六倍。为探究人类注意力建模对机器人自主性的价值,并验证该高效架构的实际应用潜力,我们将GazeLNN集成至基于强化学习训练的主动摄像头-机器人控制策略中,实现了自主导航中的类人注视引导感知,已在真实飞行机器人上成功部署并验证。
原文摘要 · Abstract (English)
Human visual attention relies on structured scanpaths to efficiently process scenes, yet instilling this behavior into robot autonomy is in its infancy and hindered by the high,computational costs of existing predictive models. To address this, we introduce GazeLNN, a computationally lightweight,scanpath prediction model that leverages Liquid Neural Networks as its recurrent engine and employs MobileNetV3 for feature extraction. Operating auto-regressively, the architecture predicts sequential fixation heatmaps conditioned on the current visual stimulus and fixation history. Despite requiring only 0.61 GFLOPs, GazeLNN achieves state-of-the-art performance on the MIT Low Resolution dataset achieving 0.47 ScanMatch score. It outperforms existing recurrent baselines across diverse evaluation metrics, while reducing computational costs by 99.40% and accelerating inference by up to six times. To investigate the role of human attention modeling in robot autonomy and demonstrate the practical utility of this highly efficient architecture, we integrate GazeLNN into an active camera-robot control policy trained via Reinforcement Learning. This integration enables human-fixation-guided perception during autonomous navigation, validated through successful real-world deployments on an aerial robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。