首个面向人形机器人在化学实验室进行精细操作的仿真与评测基准
Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

- 构建真实化学实验台的高保真仿真环境,包含30+功能部件和粉末物理模拟
- 定义6种原子操作与7步称量流程,首次实现从操作到测量的全流程评估
- 揭示任务完成度与实验精度间的差距,适合研发实验室机器人系统的研究者
实验室自动化在机器人平台与人工智能驱动的科学推理方面取得了显著进展。然而,许多实验操作(如固-固转移)仍具有高度动态性,需实时适应不同材料与实验条件。这类高精度操作难以标准化,促使采用具备灵巧手的人形机器人。尽管如此,现有研究缺乏对人形机器人在精密实验室环境中操作能力的评测基准。我们提出Labimus,据我们所知,首个面向有机化学实验室中人形机器人灵巧操作的仿真与评测基准。通过真实到仿真建模,重建了超过30个功能真实的实验室资产,覆盖常规有机化学实验的核心操作。该基准集成了带关节的实验仪器、基于粒子的粉末物理模型以及闭环仪器读数,实现了从操作到测量的完整流程。进一步依据真实标准操作规程,定义了六种原子操作和七步固体称量工作流。我们引入一种注重精度的评估协议,联合衡量任务完成率、实验精度和长时程执行表现。在程序化布局与环境扰动下,对三种代表性策略进行基准测试。结果揭示出精度缺口:成功完成任务的策略仍可能无法满足实验规程要求的定量容差。该基准暴露了任务完成与实验有效性之间的根本断层,为开发可靠人形机器人用于科学实验室提供了新测试平台。
原文摘要 · Abstract (English)
Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipulations are difficult to standardize, motivating the use of humanoid robots with dexterous hands. Despite this opportunity, no existing benchmark evaluates humanoid manipulation in precision-critical laboratory environments. We present Labimus, to our knowledge, the first benchmark for humanoid dexterous manipulation in organic chemistry laboratories. Labimus reconstructs over 30 functionally faithful assets from real organic chemistry workstations through real-to-sim modeling, collectively covering the core operations of routine organic chemistry experiments. The benchmark integrates articulated laboratory instruments, particle-based powder physics, and closed-loop instrument readouts, enabling a complete manipulation-to-measurement pipeline. It further defines six atomic operations and a seven-step solid-weighing workflow derived from real laboratory standard operating procedures. We introduce a precision-aware evaluation protocol designed to jointly measure task completion, experimental precision, and long-horizon execution. We benchmark three representative policies under procedural layouts and environmental perturbations. Results reveal a precision gap: policies that successfully complete laboratory tasks can still fail to satisfy the quantitative tolerances required by experimental protocols. Our benchmark exposes a fundamental disconnect between task completion and experimental validity, providing a new testbed for developing reliable humanoid robots for scientific laboratories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。