实测四类模拟策略对机器人抓取泛化能力的影响
Grounding Sim-to-Real Generalization in Robotic Manipulation: An Empirical Study with Vision-Language-Action Models
- 用多层级随机化、逼真渲染等四类方法测试模拟到现实的迁移效果
- 10,000次真实实验表明,物理建模精度对泛化性能影响最大
- 开源平台与评估协议,助力未来机器人控制研究标准化
学习通用机器人操控策略通常依赖大规模数据集。由于真实数据采集成本高,一个实用替代方案是通过仿真生成合成数据。然而,合成数据常与真实分布存在显著差距。尽管已有诸多研究提出算法以弥合模拟到现实的差异,但缺乏在真实操控任务中系统验证这些方法的研究,尤其是针对视觉-语言-动作(VLA)等通用策略。本研究从四个维度实证考察模拟到现实泛化的关键因素:多层级域随机化、逼真渲染、物理建模真实性及强化学习更新。为支持研究,设计了一套全面的评估协议,量化真实世界操控任务表现,涵盖背景、光照、干扰物、物体类型和空间特征等关键变化。通过超过10,000次真实世界试验,得出模拟到现实迁移的关键洞见。为推动后续研究,公开机器人平台与评估协议,建立可复现的标准基准。
原文摘要 · Abstract (English)
Learning a generalist control policy for robotic manipulation typically relies on large-scale datasets. Given the high cost of real-world data collection, a practical alternative is to generate synthetic data through simulation. However, the resulting synthetic data often exhibits a significant gap from real-world distributions. While many prior studies have proposed algorithms to bridge the Sim-to-Real discrepancy, there remains a lack of principled research that grounds these methods in real-world manipulation tasks, particularly their performance on generalist policies such as Vision-Language-Action (VLA) models. In this study, we empirically examine the primary determinants of Sim-to-Real generalization across four dimensions: multi-level domain randomization, photorealistic rendering, physics-realistic modeling, and reinforcement learning updates. To support this study, we design a comprehensive evaluation protocol to quantify the real-world performance of manipulation tasks. The protocol accounts for key variations in background, lighting, distractors, object types, and spatial features. Through experiments involving over 10k real-world trials, we derive critical insights into Sim-to-Real transfer. To inform and advance future studies, we release both the robotic platforms and the evaluation protocol for public access to facilitate independent verification, thereby establishing a realistic and standardized benchmark for robotic manipulation policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。