视觉与本体感知融合在机械臂操作中为何失效?
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
- 通过时序控制实验发现,运动转换阶段视觉作用有限
- 本体感知因收敛快而主导训练,抑制视觉学习
- 提出GAP算法动态调节梯度,实现双模态协同
本体感知信息对精确伺服控制至关重要,能提供实时机器人状态。其与视觉的协同有望提升复杂任务下的操作策略性能。然而,近期研究对视觉-本体感知策略的泛化能力存在不一致观察。本文通过时序控制实验发现,在机器人运动转换阶段(需目标定位)时,视觉模态的作用受限。进一步分析表明,策略在训练中自然倾向于收敛更快的简洁本体感知信号,从而主导优化并抑制视觉模态的学习。为此,我们提出梯度调整与阶段引导算法(GAP),通过本体感知估计轨迹中每个时间步属于运动转换阶段的概率,并在此基础上细粒度调节本体感知梯度幅度,实现视觉-本体感知策略的动态协作。大量实验表明,GAP适用于模拟与真实环境,支持单臂与双臂配置,兼容传统及视觉-语言-动作模型。该工作为机器人操作中视觉-本体感知策略的开发提供了重要洞见。
原文摘要 · Abstract (English)
Proprioceptive information is critical for precise servo control by providing real-time robotic states. Its collaboration with vision is highly expected to enhance performances of the manipulation policy in complex tasks. However, recent studies have reported inconsistent observations on the generalization of vision-proprioception policies. In this work, we investigate this by conducting temporally controlled experiments. We found that during task sub-phases that robot's motion transitions, which require target localization, the vision modality of the vision-proprioception policy plays a limited role. Further analysis reveals that the policy naturally gravitates toward concise proprioceptive signals that offer faster loss reduction when training, thereby dominating the optimization and suppressing the learning of the visual modality during motion-transition phases. To alleviate this, we propose the Gradient Adjustment with Phase-guidance (GAP) algorithm that adaptively modulates the optimization of proprioception, enabling dynamic collaboration within the vision-proprioception policy. Specifically, we leverage proprioception to capture robotic states and estimate the probability of each timestep in the trajectory belonging to motion-transition phases. During policy learning, we apply fine-grained adjustment that reduces the magnitude of proprioception's gradient based on estimated probabilities, leading to robust and generalizable vision-proprioception policies. The comprehensive experiments demonstrate GAP is applicable in both simulated and real-world environments, across one-arm and dual-arm setups, and compatible with both conventional and Vision-Language-Action models. We believe this work can offer valuable insights into the development of vision-proprioception policies in robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。