提出软体机器人视觉语言操控基准,解决变形控制与感知难题
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics

- 构建耦合弹性力约束的软体仿真环境,支持复杂接触交互
- 生成6300个场景及专家轨迹,支持策略训练与评估
- 揭示视觉感知误差是失败主因,凸显可变形性利用不足
现有视觉语言操控研究多聚焦刚性机械臂,其固定结构在杂乱或狭小空间中适应性差。软体机械臂凭借可变形性具吸引力,但面临本体感知不可靠、底层驱动分散等挑战。为此,我们提出 ManiSoft,一个面向软体臂的视觉语言操控基准。该基准包含定制仿真器,通过弹性力约束实现真实软体动力学与丰富接触交互;定义四类任务,涵盖末端执行器协同到避障等可变形控制特性。为支持策略训练与评估, ManiSoft 提供自动化流水线,生成6,300个多样化场景及对应专家轨迹。高阶规划器将任务分解为航点序列,低阶强化学习策略生成扭矩指令跟踪航点。基准测试三类典型策略模型显示,在干净场景表现尚可,但随机化条件下性能显著下降。可视化分析表明,失败主因是视觉对本体状态估计不准,且未能有效利用可变形性实现自适应避障。我们期望 ManiSoft 能成为连接刚性与软体臂在视觉语言操控中的关键测试平台。代码与数据集已开源:https://buaa-colalab.github.io/ManiSoft。
原文摘要 · Abstract (English)
Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer an appealing alternative due to their deformability, but confront challenges such as unreliable proprioception and distributed low-level actuation. To investigate these challenges, we introduce \ManiSoft, a benchmark for vision-language manipulation with soft arms. ManiSoft features a tailored simulator that couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint. On this basis, ManiSoft defines four tasks, each highlighting distinct aspects of deformable control, from basic end-effector coordination to obstacle avoidance. To support policy training and evaluation, \ManiSoft{} includes an automated pipeline that generates $6{,}300$ diverse scenes and corresponding expert trajectories. To produce high-quality trajectories at scale, we first employ a high-level planner to decompose each task into a sequence of waypoints, followed by a low-level reinforcement learning policy that generates torque commands to track waypoints. Benchmarking three representative policy models shows relatively promising results in clean scenes but substantial performance drop under randomization. Visualization analysis indicates that failures stem primarily from inaccurate visual estimation of proprioceptive state and limited exploitation of deformability for adaptive obstacle avoiding. We anticipate ManiSoft to serve as a valuable testbed, bridging the gap between rigid and soft arms in the context of vision-language manipulation. Out codes and datasets are released at https://buaa-colalab.github.io/ManiSoft.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。