arXiv:2503.14329cs.CV2025-03ICCV被引 9

让机器人像进化一样学会抓取,通过反馈不断优化策略。

EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment

  • 基于偏好对齐的迭代优化,持续改进抓取动作。
  • 在4个数据集上达成最高抓取成功率与采样效率。
  • 适合需要自适应抓取的复杂场景机器人应用。

灵巧机械手在复杂环境中泛化能力差,主要受限于低多样性训练数据。现实世界场景无限多样,难以穷尽所有可能。受生物进化启发,我们提出EvolvingGrasp,一种通过高效偏好对齐实现持续进化的抓取生成方法。核心是手部姿态偏好优化(HPO),使模型能从正负反馈中持续学习,逐步优化抓取策略。为提升在线调整的效率与可靠性,HPO引入物理一致性模型,加速推理、减少偏好微调所需步数,并保证动作物理合理性。在四个基准数据集上的大量实验表明,该方法在抓取成功率和采样效率方面均达到当前最优。结果验证了EvolvingGrasp可实现鲁棒、物理可行且偏好一致的抓取,在仿真与真实场景中均有效。

原文摘要 · Abstract (English)

Dexterous robotic hands often struggle to generalize effectively in complex environments due to the limitations of models trained on low-diversity data. However, the real world presents an inherently unbounded range of scenarios, making it impractical to account for every possible variation. A natural solution is to enable robots learning from experience in complex environments, an approach akin to evolution, where systems improve through continuous feedback, learning from both failures and successes, and iterating toward optimal performance. Motivated by this, we propose EvolvingGrasp, an evolutionary grasp generation method that continuously enhances grasping performance through efficient preference alignment. Specifically, we introduce Handpose wise Preference Optimization (HPO), which allows the model to continuously align with preferences from both positive and negative feedback while progressively refining its grasping strategies. To further enhance efficiency and reliability during online adjustments, we incorporate a Physics-aware Consistency Model within HPO, which accelerates inference, reduces the number of timesteps needed for preference finetuning, and ensures physical plausibility throughout the process. Extensive experiments across four benchmark datasets demonstrate state of the art performance of our method in grasp success rate and sampling efficiency. Our results validate that EvolvingGrasp enables evolutionary grasp generation, ensuring robust, physically feasible, and preference-aligned grasping in both simulation and real scenarios.

抓取生成进化优化偏好对齐机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。