arXiv:2605.30740cs.ROcs.AI2026-05中稿 · the 19th Internati…

提出安全通用的机械臂操作框架,提升对复杂铰链物体的适应性与操作安全性。

GSAM: A Generalizable and Safe Robotic Framework for Articulated Object Manipulation

论文配图:GSAM: A Generalizable and Safe Robotic Framework for Articulated Object Manipulation
图 1 · 摘自论文原文
  • 用视觉感知+语言模型推理修正姿态估计,融合常识判断避免误判。
  • 在50种铰链任务中成功率提升36.0%,标准差降低3.1%。
  • 适合需要高安全性和泛化能力的家用或服务机器人场景。

铰链类物体操作是服务机器人面临的独特挑战。现有方法采用端到端策略学习、视觉-运动规划及大语言/视觉语言模型(LLM/VLM),但常忽视铰链物体多样性及末端执行器与把手间的复杂交互,导致泛化能力弱且易发生破坏性碰撞。为此,我们提出GSAM——一种通用且安全的铰链物体操作框架。具体而言,基于视觉的感知器生成运动学参数;针对感知器预训练标记产生的原始估计可能违背常识的问题,提出微调的VLM-Rfiner,利用思维链(COT)常识推理进行修正。为防止破坏性碰撞,设计交互约束函数生成器,将铰链物体、交互位姿及避障知识整合为基底,再由LLM将其功能化并应用于轨迹与姿态规划。基于运动学的操纵规划器验证轨迹与姿态的可达性。在50个铰链任务、5类物体及50组随机初始化的末端-把手配置上实验表明,相比最优基线,GSAM分别将标准差降低3.1%,成功率提升36.0%,充分证明其在实际场景中的卓越泛化性与交互安全性。

原文摘要 · Abstract (English)

Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, and large-language/visual-language model (LLM/VLM), but often overlook the diversity of articulated objects and the complexity of interactions between end-effector and handle, leading to limited generalization and destructive collisions. To address this, we propose GSAM, a generalizable and safe robotic framework for articulated object manipulation. Specifically, a vision-based perceiver generates the kinematic parameters. Considering that pre-trained markers in perceiver yield raw estimations that may deviate from commonsense, we present a f ine-tuned VLM-based refiner, using chain-of-thought (COT) commonsense reasoning to refine perception. To prevent destructive collisions, we design an interaction constraint function generator, integrating articulated object, interaction pose, and obstacle avoidance knowledge into a base. LLM then functionalize these constraints and apply them to trajectory and posture planning. A kinematic-aware manipulation planner verifies reachability for trajectory and posture. Experiments on 50 hinge tasks across 5 object categories and 50 randomly initialized end-effectorhandle configurations show that GSAM reduces standard deviation by 3.1% and improves manipulation success rate by 36.0% compared to the best baseline, respectively demonstrating the superior object generalization and interaction safety of GSAM in practical scenarios.

机器人操作安全控制视觉语言模型通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。