arXiv:2606.13578cs.CLcs.AI2026-06被引 2

让AI在实验室里动手做实验,突破从写方案到执行的鸿沟。

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

论文配图:LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
图 1 · 摘自论文原文
  • 用仿真生成实验室操作数据,支持多种机器人执行
  • 分两阶段训练:先让模型理解动作,再学精准控制
  • 在真实实验场景中表现最优,适合科研自动化

科学实验日益依赖AI进行文献分析、假设生成和实验设计,但实际操作仍需人工。现有视觉-语言-动作(VLA)模型多基于家庭场景训练,难以应对实验室特有的仪器、透明液体和固定流程。本文提出数据与模型双突破:构建RoboGenesis仿真引擎,自动生成结构化实验示范;提出LabVLA模型,采用两阶段训练——先通过FAST动作标记预训练使模型具备动作感知能力,再通过流匹配微调引入DiT动作专家,在知识隔离下实现精准控制。在LabUtopia基准测试中,LabVLA在分布内与分布外设置下均达到最高平均成功率。

原文摘要 · Abstract (English)

Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks alongside model design. To address the data side, we build RoboGenesis, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles. On the policy side, we present LabVLA, trained with a two-stage recipe: FAST action token pretraining first makes the Qwen3-VL-4B-Instruct backbone action aware before any continuous control is learned, and flow matching posttraining then attaches a DiT action expert under knowledge insulation. On the LabUtopia benchmark, LabVLA achieves the highest average success rate among all evaluated baselines under both in-distribution and out-of-distribution settings.

实验室自动化视觉语言动作仿真训练多机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。