arXiv:2608.31002cs.ROcs.CV2026-08

双机械臂多视角采集数据集,助力机器人精准感知物体三维结构。

DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception

论文配图:DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception
图 1 · 摘自论文原文
  • 双机械臂对称布局,同步采集RGB、深度与红外多模态数据。
  • 224帧测试中点云到表面距离中位数仅2.13毫米,96.56%点在10毫米内。
  • 适合研究多视角重建、协同感知与基于部分观测的智能推理。

单视角机器人感知常受限于自遮挡和表面可见性不足。本文提出DARP(Dual-Arm Robotic Perception)数据集,通过置于共享桌面工作区两侧的两个独立运动眼-手机械臂,实现以物体为中心的多视角感知。每臂搭载Intel RealSense传感器,持续记录RGB、深度与立体红外数据,同步记录机器人关节状态用于位姿恢复。物体无固定姿态或标记位置,采集过程包含自动定位、跨臂确认、自适应视角生成与连续多模态记录。DARP包含10个独特桌面物体,保留原始传感器数据、机器人状态日志、物体级元数据及标定信息,支持在统一度量坐标系下重建相机轨迹。为评估几何一致性,我们构建确定性多视角融合流程,将校准后的RGB-D观测转化为互补的局部点云与测量表面网格,未使用学习或生成式补全方法。在224个保留的RGB-D关键帧(共1,563,466个三维查询点)上评估,中位点到网格距离为2.13毫米,均方根误差为4.04毫米,96.56%的点距测量表面小于10毫米。DARP旨在作为多视角重建、协同机器人感知、多模态融合、主动感知及未来基于部分观测的学习推理的可复用资源。

原文摘要 · Abstract (English)

Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Robotic Perception) https://doi.org/10.21227/rmv3-be47, a calibrated dual-arm RGB-D-IR dataset for object-centered robotic perception using two independently moving eye-in-hand manipulators positioned on opposite sides of a shared tabletop workspace. Each arm carries an Intel RealSense sensor that continuously records RGB, depth, and stereo infrared data while synchronized robot joint states are logged for pose recovery. Objects are placed without fixed poses or marked locations, and the acquisition procedure performs automatic localization, cross-arm confirmation, adaptive viewpoint generation, and continuous multimodal recording. DARP contains ten unique tabletop objects and preserves the original sensor recordings, robot-state logs, object-level metadata, and calibration information required to reconstruct camera trajectories in a shared metric frame. To evaluate the geometric consistency of the acquisition, we implement a deterministic multi-view fusion pipeline that converts calibrated RGB-D observations into complementary partial point clouds and measured surface meshes without using learned or generative completion methods. Evaluation on 224 held-out RGB-D keyframes comprising 1,563,466 three-dimensional query points yields a median point-to-mesh distance of 2.13~mm and an RMSE of 4.04~mm, with 96.56\% of points within 10~mm of the measured-surface mesh. DARP is intended as a reusable resource for multi-view reconstruction, collaborative robotic perception, multimodal fusion, active perception, and future learning-based reasoning over partial object observations.

多视角感知机器人感知数据集点云重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。