用四元数建模柔性物体运动,提升长期预测精度与物理合理性。
RopeDreamer: A Kinematic Recurrent State Space Model for Dynamics of Flexible Deformable Linear Objects

- 用四元数序列表示物体姿态,天然约束形变符合物理规律。
- 50步预测误差降低40.52%,推理速度提升31.17%。
- 适合需要长时序规划的柔性物体抓取任务。
柔性线性物体(DLO)的机器人操作因高维非线性动力学和接触任务中拓扑一致性维护困难而极具挑战。现有数据驱动方法虽使用循环神经网络和图神经网络建模动力学,但常出现自交、缠绕、链段拉伸等非物理形变。本文提出一种基于四元数运动链的递归状态空间模型,将DLO编码为相对旋转序列(四元数),而非独立笛卡尔坐标,从而在隐空间内天然约束于物理可行流形,保持链段长度恒定。进一步设计双解码器结构,分离状态重建与未来状态预测,促使隐空间捕获真实变形机制。在包含多重自交的复杂抓放轨迹大规模模拟数据集上评估,本方法在50步开环预测中误差较最优基线降低40.52%,推理时间减少31.17%。在多交叉场景中仍保持优异拓扑一致性,验证其作为长时序操作规划基本构件的有效性。
原文摘要 · Abstract (English)
The robotic manipulation of Deformable Linear Objects (DLOs) is a fundamental challenge due to the high-dimensional, non-linear dynamics of flexible structures and the complexity of maintaining topological integrity during contact-rich tasks. While recent data-driven methods have utilized Recurrent and Graph Neural Networks for dynamics modeling, they often struggle with self-intersections and non-physical deformations, such as tangling and link stretching. In this paper, we propose a latent dynamics framework that combines a Recurrent State Space Model with a Quaternionic Kinematic Chain representation to enable robust, long-term forecasting of DLO states. By encoding the DLO as a sequence of relative rotations (quaternions) rather than independent Cartesian positions, we inherently constrain the model to a physically valid manifold that preserves link-length constancy. Furthermore, we introduce a dual-decoder architecture that decouples state reconstruction from future-state prediction, forcing the latent space to capture the underlying physics of deformation. We evaluate our approach on a large-scale simulated dataset of complex pick-and-place trajectories involving self-intersections. Our results demonstrate that the proposed model achieves a 40.52% reduction in open-loop prediction error over 50-step horizons compared to the state-of-the-art baseline, while reducing inference time by 31.17%. Our model further maintains superior topological consistency in scenarios with multiple crossings, proving its efficacy as a compositional primitive for long-horizon manipulation planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。