arXiv:2601.21712cs.RO2026-01被引 4

让双臂机器人安全执行指令操作,避免自碰撞。

CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation

  • 用视觉语言动作模型结合风险估计器预测自碰撞概率。
  • 在真实机器人上测试,自碰撞减少,成功率提升。
  • 适合需要双臂协作的工业或服务机器人场景。

视觉语言动作(VLA)模型可实现指令跟随操作,但双臂部署时因未充分建模手臂与抓取物间的自碰撞问题而存在安全隐患。本文提出CoFreeVLA,通过在端到端VLA基础上增加一个短时域自碰撞风险估计算法,该算法基于本体感知、视觉嵌入和计划动作预测碰撞概率。风险估计器会屏蔽高风险指令,通过风险引导调整恢复至安全状态,并优化策略以生成更安全的执行轨迹。该估计算法先用基于模型的碰撞标签进行预训练,再在真实机器人轨迹上微调校准。在五项双臂任务中使用PiPER机械臂进行测试,CoFreeVLA相比RDT和APEX显著降低自碰撞率并提升成功率。

原文摘要 · Abstract (English)

Vision Language Action (VLA) models enable instruction following manipulation, yet dualarm deployment remains unsafe due to under modeled selfcollisions between arms and grasped objects. We introduce CoFreeVLA, which augments an endtoend VLA with a short horizon selfcollision risk estimator that predicts collision likelihood from proprioception, visual embeddings, and planned actions. The estimator gates risky commands, recovers to safe states via risk-guided adjustments, and shapes policy refinement for safer rollouts. It is pre-trained with model-based collision labels and posttrained on real robot rollouts for calibration. On five bimanual tasks with the PiPER robot arm, CoFreeVLA reduces selfcollisions and improves success rates versus RDT and APEX.

双臂控制视觉语言安全操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。