arXiv:2605.05241cs.ROcs.LG2026-05被引 1

用视觉语言大模型提升仿真到现实的机械手操作泛化能力

DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation

  • 用大模型当视觉真实度评判器,自动优化仿真参数
  • 零样本迁移下成功率达78.2%,实虚差距仅8.3%
  • 适合需要复杂抓取的机器人研究者使用

仿真到现实的迁移仍是将灵巧操作策略部署到真实机器人时的关键瓶颈。现有方法依赖人工设计的领域随机化或任务特定适应,泛化能力受限。我们提出DexSim2Real,一个集成框架,利用视觉-语言基础模型弥合灵巧操作的仿真与现实差距。系统包含三部分:(1) 基于大模型引导的领域随机化(FM-DR),通过闭环CMA-ES优化仿真参数,以视觉反馈替代纯文本指导;(2) 触觉-视觉交叉注意力策略(TVCAP),实现零样本仿真到现实强化学习;(3) 渐进式技能课程(PSC),基于LLM的任务分解与接触密集型任务难度调度。在六个挑战性操作任务上的盲评实验表明,DexSim2Real平均真实世界成功率达78.2%,优于DrEureka和DeXtreme,且仿真到现实性能差距缩小至8.3%。

原文摘要 · Abstract (English)

Sim-to-real transfer remains a critical bottleneck for deploying dexterous manipulation policies learned in simulation to real-world robots. Existing approaches rely on manually designed domain randomization or task-specific adaptation, limiting their generalizability across diverse manipulation scenarios. We present DexSim2Real, an integrated framework that leverages vision-language foundation models to bridge the sim-to-real gap for dexterous manipulation. Our system combines three components: (1) Foundation Model-Guided Domain Randomization (FM-DR), which uses a vision-language model as a visual realism critic to optimize simulation parameters via closed-loop CMA-ES, complementing text-based approaches like DrEureka with direct visual feedback; (2) a Tactile-Visual Cross-Attention Policy (TVCAP) that adapts cross-attention visuo-tactile fusion to zero-shot sim-to-real RL; and (3) a Progressive Skill Curriculum (PSC) that builds on LLM-based task decomposition with a difficulty scheduler tailored to contact-rich dexterous tasks. Extensive experiments on six challenging manipulation tasks with blinded evaluation demonstrate that DexSim2Real achieves a 78.2% average real-world success rate, outperforming DrEureka and DeXtreme while reducing the sim-to-real performance gap to only 8.3%.

灵巧操作仿真迁移大模型触觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。