构建生物实验室机器人仿真与评测平台,推动高精度科学任务自动化研究。
AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
- 通过数字化仪器与物理引擎,实现实验室设备的高保真模拟。
- 涵盖三类难度任务,验证模型在复杂实验流程中的指令理解与操作精度。
- 适合研究通用机器人系统在专业科学场景的应用,尤其关注多模态交互。
视觉-语言-动作(VLA)模型通过融合视觉、语言和本体感知信息生成动作轨迹,在家庭任务中展现潜力。然而,专业科学领域仍缺乏系统性评估。我们提出AutoBio,一个面向数字生物学实验室的仿真框架与评测基准,旨在评估机器人在生物实验环境中的自动化能力——该场景兼具结构化流程、高精度要求与多模态交互特征。AutoBio通过仪器数字化流水线、适用于实验室常见机械装置的专用物理插件,以及支持动态界面与透明材料的基于物理渲染堆栈,扩展了现有仿真能力。基准包含三类难度层级的生物学任务,可标准化评估语言引导的机器人操作性能。我们提供演示生成工具及与主流VLA模型的无缝集成接口。对两种先进VLA模型的基线测试揭示其在精确操控、视觉推理和指令遵循方面存在显著差距。通过开放AutoBio,我们希望推动通用机器人系统在复杂、高精度、多模态专业环境中的研究。仿真器与基准已公开,支持可复现研究。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have shown promise as generalist robotic policies by jointly leveraging visual, linguistic, and proprioceptive modalities to generate action trajectories. While recent benchmarks have advanced VLA research in domestic tasks, professional science-oriented domains remain underexplored. We introduce AutoBio, a simulation framework and benchmark designed to evaluate robotic automation in biology laboratory environments--an application domain that combines structured protocols with demanding precision and multimodal interaction. AutoBio extends existing simulation capabilities through a pipeline for digitizing real-world laboratory instruments, specialized physics plugins for mechanisms ubiquitous in laboratory workflows, and a rendering stack that support dynamic instrument interfaces and transparent materials through physically based rendering. Our benchmark comprises biologically grounded tasks spanning three difficulty levels, enabling standardized evaluation of language-guided robotic manipulation in experimental protocols. We provide infrastructure for demonstration generation and seamless integration with VLA models. Baseline evaluations with two SOTA VLA models reveal significant gaps in precision manipulation, visual reasoning, and instruction following in scientific workflows. By releasing AutoBio, we aim to catalyze research on generalist robotic systems for complex, high-precision, and multimodal professional environments. The simulator and benchmark are publicly available to facilitate reproducible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。