arXiv:2505.18472cs.RO2025-05被引 22

构建首个可复现的触觉视觉协同机器人操作评估基准

ManiFeel: Benchmarking and Understanding Visuotactile Manipulation Policy Learning

  • 设计多模态仿真环境,系统评测触觉+视觉策略性能
  • 实验证明触觉显著提升复杂场景下操作成功率
  • 适合研究机器人感知、具身智能与多模态学习的学者

监督式视觉运动策略在机器人操作中表现强劲,但在视觉受限场景(如狭小空间、暗光环境)或需精准感知物体属性与交互时仍面临挑战。此时触觉反馈至关重要。尽管视觉模仿学习得益于高质量仿真基准,但触觉视觉协同领域尚缺乏全面可靠的评估体系。为此,我们提出 ManiFeel——一个可复现、可扩展的仿真基准,涵盖丰富接触密集型与视觉挑战性任务,提供跨传感模态、触觉表征与策略架构的模块化评估流程,并包含真实世界验证。大量实验表明,触觉显著提升多种操作场景下的策略性能,从依赖接触的任务到视觉受限环境均有改善。结果揭示不同触觉模态的任务适应性差异,识别出鲁棒触觉视觉策略的关键设计原则与开放挑战。真实世界测试进一步验证了 ManiFeel 的可靠性和指导价值。为促进可复现研究,我们将开源代码、数据集、训练日志及预训练权重,助力通用触觉视觉策略的发展。

原文摘要 · Abstract (English)

Supervised visuomotor policies have shown strong performance in robotic manipulation but often struggle in tasks with limited visual inputs, such as operations in confined spaces and dimly lit environments, or tasks requiring precise perception of object properties and environmental interactions. In such cases, tactile feedback becomes essential for manipulation. While the rapid progress of supervised visuomotor policies has benefited greatly from high-quality, reproducible simulation benchmarks in visual imitation, the visuotactile domain still lacks a similarly comprehensive and reliable benchmark for large-scale and rigorous evaluation. To address this, we introduce ManiFeel, a reproducible and scalable simulation benchmark designed to systematically study supervised visuotactile policy learning. ManiFeel offers a diverse suite of contact-rich and visually challenging manipulation tasks, a modular evaluation pipeline spanning sensing modalities, tactile representations, and policy architectures, as well as real-world validation. Through extensive experiments, ManiFeel demonstrates how tactile sensing enhances policy performance across diverse manipulation scenarios, ranging from precise contact-driven operations to visually constrained settings. In addition, the results reveal task-dependent strengths of different tactile modalities and identify key design principles and open challenges for robust visuotactile policy learning. Real-world evaluations further confirm that ManiFeel provides a reliable and meaningful foundation for benchmarking and future visuotactile policy development. To foster reproducibility and future research, we will release our codebase, datasets, training logs, and pretrained checkpoints, aiming to accelerate progress toward generalizable visuotactile policy learning and manipulation.

机器人操作触觉感知多模态学习仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。