arXiv:2512.17764cs.RO2025-12被引 1

用扩散模型实现被遮挡下柔性线状物的精准状态估计与跟踪

UniStateDLO: Unified Generative State Estimation and Tracking of Deformable Linear Objects Under Occlusion for Constrained Manipulation

  • 基于扩散模型将局部点云映射为高维状态,统一处理单帧估计与跨帧跟踪
  • 在严重遮挡下仍能实时生成全局平滑、局部精确的状态预测,优于所有现有方法
  • 仅用合成数据训练即可零样本迁移到真实场景,适合复杂空间中的柔性物体操作

柔性线状物(如电缆、绳索、电线)的感知是成功执行下游操作的基础。尽管基于视觉的方法已广泛研究,但在受限操作环境中,由于周围障碍物、大而变化的形变及视角有限,遮挡问题普遍存在,导致感知性能显著下降。此外,状态空间维度高、视觉特征不明显、传感器噪声等因素进一步加剧了可靠感知的难度。为此,本文提出UniStateDLO,首个采用深度学习的完整柔性线状物感知流水线,可在严重遮挡条件下实现鲁棒性能,涵盖单帧状态估计与跨帧状态跟踪,输入为部分点云。两项任务均建模为条件生成问题,利用扩散模型强大的非线性映射能力,从高度局部观测中恢复高维状态。UniStateDLO有效应对初始遮挡、自遮挡及多物体遮挡等多种模式。其训练仅依赖大规模合成数据集,实现零样本仿真到现实迁移,无需真实世界数据。大量仿真与真实实验表明,UniStateDLO在估计与跟踪任务上全面超越现有最优方法,实现实时、全局平滑且局部精确的状态预测,即使在强遮挡下亦表现优异。将其作为闭环柔性线状物操作系统的前端模块,验证了其在复杂三维受限环境中的稳定反馈控制能力。

原文摘要 · Abstract (English)

Perception of deformable linear objects (DLOs), such as cables, ropes, and wires, is the cornerstone for successful downstream manipulation. Although vision-based methods have been extensively explored, they remain highly vulnerable to occlusions that commonly arise in constrained manipulation environments due to surrounding obstacles, large and varying deformations, and limited viewpoints. Moreover, the high dimensionality of the state space, the lack of distinctive visual features, and the presence of sensor noises further compound the challenges of reliable DLO perception. To address these open issues, this paper presents UniStateDLO, the first complete DLO perception pipeline with deep-learning methods that achieves robust performance under severe occlusion, covering both single-frame state estimation and cross-frame state tracking from partial point clouds. Both tasks are formulated as conditional generative problems, leveraging the strong capability of diffusion models to capture the complex mapping between highly partial observations and high-dimensional DLO states. UniStateDLO effectively handles a wide range of occlusion patterns, including initial occlusion, self-occlusion, and occlusion caused by multiple objects. In addition, it exhibits strong data efficiency as the entire network is trained solely on a large-scale synthetic dataset, enabling zero-shot sim-to-real generalization without any real-world training data. Comprehensive simulation and real-world experiments demonstrate that UniStateDLO outperforms all state-of-the-art baselines in both estimation and tracking, producing globally smooth yet locally precise DLO state predictions in real time, even under substantial occlusions. Its integration as the front-end module in a closed-loop DLO manipulation system further demonstrates its ability to support stable feedback control in complex, constrained 3-D environments.

柔性物体状态估计扩散模型遮挡处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。