用视觉建模连续体机器人形状,实现可解释的几何感知控制。
Shape-Interpretable Visual Self-Modeling Enables Geometry-Aware Continuum Robot Control
- 通过贝塞尔曲线编码多视角图像,构建可解释的三维形状空间。
- 神经微分方程自学习形态与末端运动,误差分别低于1.56%和2%。
- 支持避障与自运动,适合需要几何感知的复杂环境操作。
连续体机器人具有高灵活性和冗余性,适用于复杂环境中的安全交互,但其连续变形和非线性动力学给感知、建模与控制带来根本挑战。现有基于视觉的控制方法多依赖端到端学习,虽能实现形状调节,但缺乏对机器人几何结构或与环境交互的显式认知。本文提出一种形状可解释的视觉自建模框架,利用多视角平面图像,通过贝塞尔曲线表示将机器人形状编码为紧凑且物理意义明确的形状空间,唯一表征其三维构型。基于该表示,采用神经常微分方程直接从数据中自学习形态与末端执行器动力学,实现无需解析模型或密集身体标记的混合形状-位置控制。所学形状空间的显式几何结构使机器人能够推理自身与周围环境的关系,支持障碍物避让和自运动等环境感知行为,同时保持末端执行器目标。在缆绳驱动的连续体机器人上实验表明,形状误差小于图像分辨率的1.56%,末端误差小于机器人长度的2%,并在受限环境中表现出鲁棒性能。本工作将视觉形状表示从二维观测提升为可解释的三维自模型,为基于视觉的端到端控制提供了原理性替代方案,推动了连续体机器人的自主几何感知操控发展。
原文摘要 · Abstract (English)
Continuum robots possess high flexibility and redundancy, making them well suited for safe interaction in complex environments, yet their continuous deformation and nonlinear dynamics pose fundamental challenges to perception, modeling, and control. Existing vision-based control approaches often rely on end-to-end learning, achieving shape regulation without explicit awareness of robot geometry or its interaction with the environment. Here, we introduce a shape-interpretable visual self-modeling framework for continuum robots that enables geometry-aware control. Robot shapes are encoded from multi-view planar images using a Bezier-curve representation, transforming visual observations into a compact and physically meaningful shape space that uniquely characterizes the robot's three-dimensional configuration. Based on this representation, neural ordinary differential equations are employed to self-model both shape and end-effector dynamics directly from data, enabling hybrid shape-position control without analytical models or dense body markers. The explicit geometric structure of the learned shape space allows the robot to reason about its body and surroundings, supporting environment-aware behaviors such as obstacle avoidance and self-motion while maintaining end-effector objectives. Experiments on a cable-driven continuum robot demonstrate accurate shape-position regulation and tracking, with shape errors within 1.56% of image resolution and end-effector errors within 2% of robot length, as well as robust performance in constrained environments. By elevating visual shape representations from two-dimensional observations to an interpretable three-dimensional self-model, this work establishes a principled alternative to vision-based end-to-end control and advances autonomous, geometry-aware manipulation for continuum robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。