arXiv:2411.14358cs.ROcs.CV2024-11被引 1

为视障者导航打造真实室内人群环境的视觉惯性数据集

InCrowd-VI: A Realistic Visual-Inertial Dataset for Evaluating SLAM in Indoor Pedestrian-Rich Spaces for Human Navigation

  • 用Meta Aria眼镜采集58段真实人群场景,含双目图像与IMU数据
  • 地面真值轨迹精度约2厘米,但主流SLAM系统定位误差超0.5米
  • 适合研究视障导航、复杂环境下的鲁棒SLAM算法开发者使用

同步定位与地图构建(SLAM)技术可用于视障者导航,但缺乏真实人群密集环境的数据集限制了鲁棒SLAM方案的发展。为此,我们推出InCrowd-VI,一个专为室内人群密集空间中人类导航设计的视觉惯性数据集。该数据集通过Meta Aria Project眼镜录制,包含58个序列,总轨迹长度达5公里,记录时长1.5小时,涵盖RGB、立体图像和惯性测量单元(IMU)数据。数据集真实还原了行人遮挡、密度变化、复杂布局和光照波动等挑战。地面真值轨迹精度约为2厘米,来源于Meta Aria项目机器感知SLAM服务。每个序列还提供半稠密3D点云。在InCrowd-VI上对主流视觉里程计(VO)和SLAM算法的评估显示,这些系统在真实场景下表现严重受限:在挑战性条件下,定位误差超过0.5米,漂移超出1%阈值,传统方法漂移高达5-10米。尽管基于深度学习的方法保持了超过90%的姿态估计覆盖率,却无法实现实时处理速度以满足步行导航需求。这些结果凸显了建立新数据集以推动视障导航复杂室内环境研究的必要性。数据集及相关工具已公开发布于https://incrowd-vi.cloudlab.zhaw.ch/。

原文摘要 · Abstract (English)

Simultaneous localization and mapping (SLAM) techniques can be used to navigate the visually impaired, but the development of robust SLAM solutions for crowded spaces is limited by the lack of realistic datasets. To address this, we introduce InCrowd-VI, a novel visual-inertial dataset specifically designed for human navigation in indoor pedestrian-rich environments. Recorded using Meta Aria Project glasses, it captures realistic scenarios without environmental control. InCrowd-VI features 58 sequences totaling a 5 km trajectory length and 1.5 hours of recording time, including RGB, stereo images, and IMU measurements. The dataset captures important challenges such as pedestrian occlusions, varying crowd densities, complex layouts, and lighting changes. Ground-truth trajectories, accurate to approximately 2 cm, are provided in the dataset, originating from the Meta Aria project machine perception SLAM service. In addition, a semi-dense 3D point cloud of scenes is provided for each sequence. The evaluation of state-of-the-art visual odometry (VO) and SLAM algorithms on InCrowd-VI revealed severe performance limitations in these realistic scenarios. Under challenging conditions, systems exceeded the required localization accuracy of 0.5 meters and the 1\% drift threshold, with classical methods showing drift up to 5-10\%. While deep learning-based approaches maintained high pose estimation coverage (>90\%), they failed to achieve real-time processing speeds necessary for walking pace navigation. These results demonstrate the need and value of a new dataset to advance SLAM research for visually impaired navigation in complex indoor environments. The dataset and associated tools are publicly available at https://incrowd-vi.cloudlab.zhaw.ch/.

SLAM视障导航视觉惯性人群环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。