arXiv:2511.22950cs.CVcs.RO2025-11被引 2

提出RobotSeg模型与数据集,实现机器人图像视频的自动精准分割。

RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video

  • 基于SAM2改进,引入结构记忆关联器与机器人提示生成器
  • 在2.8万帧视频上达到顶尖分割性能,支持全自动标注
  • 适合机器人感知、仿真迁移与人机共融安全监控场景

准确的机器人分割是机器人感知的基础能力,可支持视觉伺服、数据增强、真实到仿真迁移及动态人机环境中的安全监控。尽管现代分割模型能力强大,但机器人分割仍具挑战,源于形态多样性、外观模糊性、结构复杂性及快速形变。为此,我们提出RobotSeg,一种面向图像与视频的机器人分割基础模型。该模型基于通用的SAM2,针对其在机器人分割中的三大局限——缺乏对关节机器人的适应性、依赖人工提示、需逐帧标注掩码——引入结构增强的记忆关联模块、机器人提示生成器与标签高效训练策略。这些创新共同实现结构感知、全自动且低标注成本的分割方案。我们进一步构建了包含超过2.8k段视频(138,000帧)的视频机器人分割(VRS)数据集,涵盖多样机器人形态与环境。大量实验表明,RobotSeg在图像与视频任务上均达到当前最优表现,为机器人感知的未来发展奠定坚实基础。

原文摘要 · Abstract (English)

Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accurate real-to-sim transfer, and reliable safety monitoring in dynamic human-robot environments. Despite the strong capabilities of modern segmentation models, surprisingly it remains challenging to segment robots. This is due to robot embodiment diversity, appearance ambiguity, structural complexity, and rapid shape changes. Embracing these challenges, we introduce RobotSeg, a foundation model for robot segmentation in image and video. RobotSeg is built upon the versatile SAM 2 foundation model but addresses its three limitations for robot segmentation, namely the lack of adaptation to articulated robots, reliance on manual prompts, and the need for per-frame training mask annotations, by introducing a structure-enhanced memory associator, a robot prompt generator, and a label-efficient training strategy. These innovations collectively enable a structure-aware, automatic, and label-efficient solution. We further construct the video robot segmentation (VRS) dataset comprising over 2.8k videos (138k frames) with diverse robot embodiments and environments. Extensive experiments demonstrate that RobotSeg achieves state-of-the-art performance on both images and videos, establishing a strong foundation for future advances in robot perception.

机器人分割视频分割自监督学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。