构建高保真实验室仿真与分级评测体系,推动科学智能体发展。
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
- 构建多物理场、化学意义明确的高保真仿真环境
- 涵盖30项任务、200+场景与仪器资产,支持长时序操作评估
- 适合研究具身智能在复杂科学实验中的泛化能力
科学具身智能体在现代实验室中承担自动化复杂实验流程的关键角色。相较于日常家居环境,实验室对物理-化学变化感知与长时程规划要求更高,是推进具身智能的理想测试平台。然而,其发展长期受限于缺乏合适的模拟器与基准。本文提出LabUtopia,一个综合性仿真与评测套件,旨在促进通用、具备推理能力的实验室具身智能体开发。它包含:i)LabSim,支持多物理场与化学意义交互的高保真模拟器;ii)LabScene,可扩展的程序化生成器,用于生成多样化科学场景;iii)LabBench,覆盖五个层级复杂度的分层基准,从原子动作到长时程移动操作。该平台支持30项不同任务,包含200多个场景与仪器资产,可在高复杂度环境中实现大规模训练与严谨评估。实验证明,LabUtopia为融合感知、规划与控制的科学智能体提供了强大平台,并为未来研究具身智能的实际能力与泛化极限提供了严格测试基准。
原文摘要 · Abstract (English)
Scientific embodied agents play a crucial role in modern laboratories by automating complex experimental workflows. Compared to typical household environments, laboratory settings impose significantly higher demands on perception of physical-chemical transformations and long-horizon planning, making them an ideal testbed for advancing embodied intelligence. However, its development has been long hampered by the lack of suitable simulator and benchmarks. In this paper, we address this gap by introducing LabUtopia, a comprehensive simulation and benchmarking suite designed to facilitate the development of generalizable, reasoning-capable embodied agents in laboratory settings. Specifically, it integrates i) LabSim, a high-fidelity simulator supporting multi-physics and chemically meaningful interactions; ii) LabScene, a scalable procedural generator for diverse scientific scenes; and iii) LabBench, a hierarchical benchmark spanning five levels of complexity from atomic actions to long-horizon mobile manipulation. LabUtopia supports 30 distinct tasks and includes more than 200 scene and instrument assets, enabling large-scale training and principled evaluation in high-complexity environments. We demonstrate that LabUtopia offers a powerful platform for advancing the integration of perception, planning, and control in scientific-purpose agents and provides a rigorous testbed for exploring the practical capabilities and generalization limits of embodied intelligence in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。