发现人形机器人双手操作中初始姿态影响手部选择,提出改进训练数据覆盖提升鲁棒性。
Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence

- 通过量化手部偏好和响应度,揭示策略诱导的初始姿态依赖问题。
- 同一初始姿态下不同策略成功率差异大,单一策略在不同姿态上表现波动显著。
- 扩充训练数据中的初始姿态覆盖可有效提升任务鲁棒性,尤其针对低表现配置。
视觉-语言-动作(VLA)策略应能在机器人初始构型变化下保持稳健,但整体任务成功率可能掩盖特定姿态下的失败和不当的手部选择。本文研究基于VLA的人形双臂操作中的初始姿态依赖性。我们将早期手部偏好定义为策略诱导的手部先验,并使用HandPriorScore、残差手部偏差和目标响应度进行量化。在多个策略和17种初始构型上的评估显示,初始姿态与策略间存在强烈交互:相同姿态在不同策略下成功率达数倍差异,而单个策略在不同姿态上性能波动显著。特定初始手臂配置可抑制或诱发不对称手部偏好,其影响方向与强度随策略而异。腕部摄像头观测也影响手部选择与任务表现。扩大训练数据中初始姿态的覆盖范围显著提升鲁棒性,而在低表现配置附近进行针对性增强可提升其成功率。对比不同训练配置发现,充分暴露于目标任务仿真环境有益,而真实或辅助数据的影响取决于姿态覆盖、仿真比例和观测可用性。本研究揭示了姿态条件化的手部先验,识别出局部初始手臂配置作为调控手部选择行为的因果控制点,并展示了数据覆盖与训练构成对初始姿态鲁棒性的影响。
原文摘要 · Abstract (English)
Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。