arXiv:2410.08172cs.ROcs.AI2024-10被引 1

提出一套评估生成式机器人仿真任务的综合框架。

On the Evaluation of Generative Robotic Simulations

  • 从质量、多样性、泛化性三方面构建评估体系。
  • 生成任务在真实性和轨迹完整性上表现良好,但泛化能力普遍不足。
  • 适合关注仿真任务质量评估的研究者和开发者。

由于获取大量真实世界数据困难,机器人仿真已成为并行训练和模拟到现实迁移的关键,凸显了可扩展仿真机器人任务的重要性。基础模型在自主生成可行机器人任务方面展现出惊人能力。然而,这一新范式也带来了对自动生成任务进行充分评估的挑战。为此,我们提出一个针对生成式仿真的综合性评估框架。该框架将评估分为三个核心方面:质量、多样性和泛化性。对于单任务质量,我们利用大语言模型和视觉-语言模型评估生成任务的真实性和生成轨迹的完整性。在多样性方面,通过任务描述的文本相似度以及基于收集任务轨迹训练的世界模型损失来衡量任务和数据多样性。在任务级泛化性方面,评估在多个生成任务上训练的策略在未见任务上的零样本泛化能力。在三个代表性任务生成流水线上的实验表明,本框架的结果与人工评估高度一致,验证了方法的可行性与有效性。研究发现,虽然某些方法可实现质量和多样性指标,但尚无单一方法在所有指标上均占优,提示需更关注多指标间的平衡。此外,分析进一步揭示当前工作普遍存在的泛化能力低下问题。

原文摘要 · Abstract (English)

Due to the difficulty of acquiring extensive real-world data, robot simulation has become crucial for parallel training and sim-to-real transfer, highlighting the importance of scalable simulated robotic tasks. Foundation models have demonstrated impressive capacities in autonomously generating feasible robotic tasks. However, this new paradigm underscores the challenge of adequately evaluating these autonomously generated tasks. To address this, we propose a comprehensive evaluation framework tailored to generative simulations. Our framework segments evaluation into three core aspects: quality, diversity, and generalization. For single-task quality, we evaluate the realism of the generated task and the completeness of the generated trajectories using large language models and vision-language models. In terms of diversity, we measure both task and data diversity through text similarity of task descriptions and world model loss trained on collected task trajectories. For task-level generalization, we assess the zero-shot generalization ability on unseen tasks of a policy trained with multiple generated tasks. Experiments conducted on three representative task generation pipelines demonstrate that the results from our framework are highly consistent with human evaluations, confirming the feasibility and validity of our approach. The findings reveal that while metrics of quality and diversity can be achieved through certain methods, no single approach excels across all metrics, suggesting a need for greater focus on balancing these different metrics. Additionally, our analysis further highlights the common challenge of low generalization capability faced by current works. Our anonymous website: https://sites.google.com/view/evaltasks.

机器人仿真生成评估任务生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。