arXiv:2603.06846cs.CVcs.RO2026-03

提出运动级刚体分割新范式,提升机器人对物理交互的理解能力

MotionBits: Video Segmentation through Motion-Level Analysis of Rigid Bodies

  • 基于运动学空间旋转变换等价性定义最小运动单元MotionBit
  • 在MoRiBo数据集上实现37.3%的mIoU提升,超越现有方法
  • 无需训练的图算法,适用于真实场景下的机器人推理与操作

刚体是现实世界中可操控的最小单元,理解其物理交互是具身推理与机器人操作的基础。准确检测、分割和追踪运动刚体对于让推理模块在多样化环境中理解并行动至关重要。然而,当前依赖语义分组的分割模型在提供任务所需的交互级线索方面能力有限。为弥补这一差距,我们提出MotionBit,一种新概念:不同于以往方法,它通过运动学空间旋转变换等价性定义运动分割的最小单元,与语义无关。本文贡献包括:(1) 提出MotionBit的概念与定义;(2) 构建手标注基准MoRiBo,用于评估机器人操作与野外人类视频中的刚体分割表现;(3) 设计无需训练的基于图的MotionBits分割方法,在MoRiBo基准上以37.3%的宏平均mIoU显著优于当前最优方法。最后,我们验证了MotionBits在下游具身推理与操作任务中的有效性,凸显其作为理解物理交互基本单元的重要性。

原文摘要 · Abstract (English)

Rigid bodies constitute the smallest manipulable elements in the real world, and understanding how they physically interact is fundamental to embodied reasoning and robotic manipulation. Thus, accurate detection, segmentation, and tracking of moving rigid bodies is essential for enabling reasoning modules to interpret and act in diverse environments. However, current segmentation models trained on semantic grouping are limited in their ability to provide meaningful interaction-level cues for completing embodied tasks. To address this gap, we introduce MotionBit, a novel concept that, unlike prior formulations, defines the smallest unit in motion-based segmentation through kinematic spatial twist equivalence, independent of semantics. In this paper, we contribute (1) the MotionBit concept and definition, (2) a hand-labeled benchmark, called MoRiBo, for evaluating moving rigid-body segmentation across robotic manipulation and human-in-the-wild videos, and (3) a learning-free graph-based MotionBits segmentation method that outperforms state-of-the-art embodied perception methods by 37.3\% in macro-averaged mIoU on the MoRiBo benchmark. Finally, we demonstrate the effectiveness of MotionBits segmentation for downstream embodied reasoning and manipulation tasks, highlighting its importance as a fundamental primitive for understanding physical interactions.

视频分割刚体运动具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。