arXiv:2410.07078cs.RO2024-10CoRL被引 11

用历史感知扩散模型解决关节物体操作中的视觉模糊问题

FlowBotHD: History-Aware Diffuser Handling Ambiguities in Articulated Objects Manipulation

  • 引入历史感知扩散网络,建模关节物体的多模态操作模式
  • 在遮挡和对称场景下,显著提升操作成功率与稳定性
  • 适合机器人抓取、自动驾驶等需处理模糊视觉信息的场景

我们提出一种新方法,用于操纵视觉上存在歧义的关节物体(如对称门或严重遮挡的门)。这类物体因缺乏区分特征(如门平面被视角遮挡)或操作方向/位置不确定(如推拉方向不明),导致操作模式存在多重可能。为此,我们设计了一种历史感知扩散网络,能够建模关节物体的多模态操作分布,并利用观测历史来区分不同模式,实现遮挡下的稳定预测。实验表明,该方法在关节物体操作任务中达到当前最优性能,尤其在存在视觉模糊的场景下表现大幅提升。

原文摘要 · Abstract (English)

We introduce a novel approach for manipulating articulated objects which are visually ambiguous, such doors which are symmetric or which are heavily occluded. These ambiguities can cause uncertainty over different possible articulation modes: for instance, when the articulation direction (e.g. push, pull, slide) or location (e.g. left side, right side) of a fully closed door are uncertain, or when distinguishing features like the plane of the door are occluded due to the viewing angle. To tackle these challenges, we propose a history-aware diffusion network that can model multi-modal distributions over articulation modes for articulated objects; our method further uses observation history to distinguish between modes and make stable predictions under occlusions. Experiments and analysis demonstrate that our method achieves state-of-art performance on articulated object manipulation and dramatically improves performance for articulated objects containing visual ambiguities. Our project website is available at https://flowbothd.github.io/.

机器人操作扩散模型视觉模糊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。