arXiv:2505.11282cs.CV2025-05CVPR被引 5

首个面向高速机器人感知的事件相机多任务数据集,支持6D位姿与动目标检测。

MTevent: A Multi-Task Event Camera Dataset for 6D Pose Estimation and Moving Object Detection

  • 构建双目事件相机+RGB相机系统,采集75个高动态场景
  • 涵盖16类物体,平均时长16秒,支持远距离、极端视角感知
  • 首次融合高速运动、长距感知与真实交互,适合高精度视觉研究

移动机器人正达到前所未有的速度,如Unitree B2和Fraunhofer O3dyn平台最高速度可达5至10米/秒。然而,由于传统RGB相机存在运动模糊且无法实时响应,高效利用高速仍具挑战。事件相机具有异步运作与低延迟传感特性,为高速机器人感知提供了有前景的替代方案。本文提出MTevent数据集,专为高度动态环境下的6D位姿估计与动目标检测设计。数据集包含双目事件相机与一个RGB相机,共采集75个场景,每个场景平均持续16秒,包含16种不同物体,在极端视角、光照变化及遮挡等挑战性条件下运行。MTevent是首个同时具备高速运动、长距离感知与真实物体交互的数据集,可显著推动事件视觉在机器人领域的研究。作为基线,我们在RGB图像上使用NVIDIA FoundationPose进行6D位姿估计,仅取得0.22的平均召回率(使用真值掩码),凸显了基于RGB方法在该场景下的局限性。数据集已开放下载:https://huggingface.co/datasets/anas-gouda/MTevent。

原文摘要 · Abstract (English)

Mobile robots are reaching unprecedented speeds, with platforms like Unitree B2, and Fraunhofer O3dyn achieving maximum speeds between 5 and 10 m/s. However, effectively utilizing such speeds remains a challenge due to the limitations of RGB cameras, which suffer from motion blur and fail to provide real-time responsiveness. Event cameras, with their asynchronous operation, and low-latency sensing, offer a promising alternative for high-speed robotic perception. In this work, we introduce MTevent, a dataset designed for 6D pose estimation and moving object detection in highly dynamic environments with large detection distances. Our setup consists of a stereo-event camera and an RGB camera, capturing 75 scenes, each on average 16 seconds, and featuring 16 unique objects under challenging conditions such as extreme viewing angles, varying lighting, and occlusions. MTevent is the first dataset to combine high-speed motion, long-range perception, and real-world object interactions, making it a valuable resource for advancing event-based vision in robotics. To establish a baseline, we evaluate the task of 6D pose estimation using NVIDIA's FoundationPose on RGB images, achieving an Average Recall of 0.22 with ground-truth masks, highlighting the limitations of RGB-based approaches in such dynamic settings. With MTevent, we provide a novel resource to improve perception models and foster further research in high-speed robotic vision. The dataset is available for download https://huggingface.co/datasets/anas-gouda/MTevent

事件相机6D位姿动态感知机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。