arXiv:2603.14397cs.RO2026-03

用事件相机提升暗光环境下机器人导航的仿真实现

eNavi: Event-based Imitation Policies for Low-Light Indoor Mobile Robot Navigation

  • 融合事件与RGB图像,通过双流编码+注意力融合建模
  • 在12种训练场景下,融合模型在暗光中误差降低37%
  • 适合做低光环境机器人控制研究的团队使用

事件相机具有高动态范围和微秒级时间分辨率,适用于快速运动或低光照条件下的室内机器人导航,而传统RGB相机在此类场景下性能下降。尽管事件感知在检测、SLAM和位姿估计方面取得进展,但针对事件流异步特性设计端到端控制策略的研究仍较少。为此,我们构建了一个基于TurtleBot 2的真实世界室内人跟随数据集,包含同步的原始事件流、RGB帧及专家控制动作,覆盖多个室内地图、正常与低光条件下的轨迹。我们还开发了多模态预处理流水线,对齐事件与RGB观测,并利用里程计重建真实动作以支持高质量模仿学习。基于该数据集,我们提出一种后融合的RGB-事件导航策略,采用双MobileNet编码器与Transformer融合模块,通过行为克隆进行训练。在12种训练变体(从单路径模仿到多路径泛化)上的系统评估表明,引入事件数据的策略,尤其是融合模型,在未见低光条件下显著提升鲁棒性,动作预测误差降低37%,而纯RGB模型则失效。数据集、同步流程与训练模型已公开:https://eventbasedvision.github.io/eNavi/

原文摘要 · Abstract (English)

Event cameras provide high dynamic range and microsecond-level temporal resolution, making them well-suited for indoor robot navigation, where conventional RGB cameras degrade under fast motion or low-light conditions. Despite advances in event-based perception spanning detection, SLAM, and pose estimation, there remains limited research on end-to-end control policies that exploit the asynchronous nature of event streams. To address this gap, we introduce a real-world indoor person-following dataset collected using a TurtleBot 2 robot, featuring synchronized raw event streams, RGB frames, and expert control actions across multiple indoor maps, trajectories under both normal and low-light conditions. We further build a multimodal data preprocessing pipeline that temporally aligns event and RGB observations while reconstructing ground-truth actions from odometry to support high-quality imitation learning. Building on this dataset, we propose a late-fusion RGB-Event navigation policy that combines dual MobileNet encoders with a transformer-based fusion module trained via behavioral cloning. A systematic evaluation of RGB-only, Event-only, and RGB-Event fusion models across 12 training variations ranging from single-path imitation to general multi-path imitation shows that policies incorporating event data, particularly the fusion model, achieve improved robustness and lower action prediction error, especially in unseen low-light conditions where RGB-only models fail. We release the dataset, synchronization pipeline, and trained models at https://eventbasedvision.github.io/eNavi/

事件相机机器人导航低光环境模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。