仅用单眼摄像头实现机器人对物体功能部件的实时语义重建与抓取。
RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera

- 通过视觉语言模型发现功能部件,异步融合几何与语义信息。
- 83.4%部分交并比,24次实测中100%抓准目标部件,21次任务成功。
- 无需CAD模型,适合复杂场景下任务驱动的机器人操作。
机器人操作需要感知系统识别把手、边缘、扳机等可操作部件,而不仅仅是物体类别或点云。本文提出RoboSeg,一种基于单眼相机的部件级语义重建系统,结合视觉语言模型(VLM)的功能部件发现、异步在线RGB-D语义重建以及任务导向的抓取生成,无需依赖CAD模型或预扫描网格。RoboSeg在初始RGB图像上查询VLM获取紧凑的功能部件提示,随后通过两条异步流进行扫描:高频几何流用于RGB-D位姿估计和截断有符号距离函数(TSDF)融合;关键帧触发的语义流用于SAM3部件掩码生成。投影后的掩码通过体素级时间投票融合为持久的部件标注点云;该地图用于将AnyGrasp 6-DoF抓取候选分配至语义部件,并选择与任务相关部件一致的抓取。RoboSeg在人工标注物体上达到83.4%的平均部分交并比(mIoU);在四个物体、八项任务的24次物理试验中,所选抓取均准确接触目标部件,21/24次任务成功。结果表明,RoboSeg可作为任务条件化操作的语义索引层,同时保留AnyGrasp作为候选生成器。
原文摘要 · Abstract (English)
Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categories or point clouds. This paper presents RoboSeg, a part-level semantic reconstructionsystem that links vision-language model (VLM) functional-partdiscovery, asynchronous online RGB-D semantic reconstruc-tion, and task-oriented grasp generation without requiring CAD models or pre-scanned meshes. RoboSeg queries a VLM onthe initial RGB observation to obtain compact functional part prompts, then scans with two asynchronous streams: a high-frequency geometry thread for RGB-D odometry and truncated signed distance function (TSDF) fusion, and a keyframe-triggered semantic thread for SAM3 part masks. Projectedmasks are fused by voxel-level temporal voting into a persistentpart-labeled point cloud; RoboSeg uses this map to assign AnyGrasp 6-DoF candidates to semantic parts and select grasps consistent with the task-relevant part label. RoboSeg reaches 83.4% mean part intersection-over-union (mIoU) over manually labeled objects; in a 24-trial physical pilot across fourobjects and eight tasks, the selected grasp contacts the requestedpart in all trials and achieves 21/24 combined task successes.These results characterize RoboSeg as a semantic indexing layerfor task-conditioned manipulation, with AnyGrasp retained asthe proposal generator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。