arXiv:2603.28422cs.RO2026-03

仅用主动双目相机,比复杂多传感器更高效地教会人形机器人操作。

Active Stereo-Camera Outperforms Multi-Sensor Setup in ACT Imitation Learning for Humanoid Manipulation

  • 用主动双目相机替代多传感器,简化感知系统。
  • 在数据有限时,双目相机达94.4%操作成功率,优于多传感器组合。
  • 触觉传感器在小数据下反而降低性能,适合大数据场景。

教人形机器人新任务的复杂性是其在工业中广泛应用的主要障碍。尽管模仿学习(IL)尤其是基于Transformer的动作分块(ACT)能实现快速任务获取,但对操纵任务最优传感硬件仍无共识。本文在配备三指手的Unitree G1人形机器人上,对14种传感器组合进行了基准测试,评估了触觉与本体感觉模态与主动视觉的融合效果。通过开发开源统一消融框架,在综合主数据集上进行传感器掩码实验,结果表明:在数据受限条件下,额外模态常导致性能下降,而最小化主动双目相机配置在空间泛化任务中取得87.5%成功,在结构化操作任务中达94.4%。相反,添加压力传感器使后一任务成功率降至67.3%,因信噪比过低。结论是:数据有限时,主动视觉在鲁棒性与复杂度间提供更优权衡;触觉模态需更大数据量才有效,因此战略性的传感器选择对高效学习至关重要。

原文摘要 · Abstract (English)

The complexity of teaching humanoid robots new tasks is one of the major reasons hindering their widespread adoption in the industry. While Imitation Learning (IL), particularly Action Chunking with Transformers (ACT), enables rapid task acquisition, there is no consensus yet on the optimal sensory hardware required for manipulation tasks. This paper benchmarks 14 sensor combinations on the Unitree G1 humanoid robot equipped with three-finger hands for two manipulation tasks. We explicitly evaluate the integration of tactile and proprioceptive modalities alongside active vision. Our analysis demonstrates that strategic sensor selection can outperform complex configurations in data-limited regimes while reducing computational overhead. We develop an open-source Unified Ablation Framework that utilizes sensor masking on a comprehensive master dataset. Results indicate that additional modalities often degrade performance for IL with limited data. A minimal active stereo-camera setup outperformed complex multi-sensor configurations, achieving 87.5% success in a spatial generalization task and 94.4% in a structured manipulation task. Conversely, adding pressure sensors to this setup reduced success to 67.3% in the latter task due to a low signal-to-noise ratio. We conclude that in data-limited regimes, active vision offers a superior trade-off between robustness and complexity. While tactile modalities may require larger datasets to be effective, our findings validate that strategic sensor selection is critical for designing an efficient learning process.

人形机器人模仿学习主动视觉传感器优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。