用协议驱动的视觉增强系统,让低成本机器人可靠完成生物实验操作。
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation

- 以实验协议为输入,分步解析并验证任务状态,实现闭环执行。
- 在15项基础任务和6种复合流程中,精准放置与透明容器操作成功率提升显著。
- 适合需要可重复、高鲁棒性生物实验自动化的研究团队使用。
生物实验室自动化可减少重复性人工工作并提升可重现性,但湿实验环境中可靠的任务执行仍具挑战。实验协议常无结构化,器皿常透明或反光,多步骤操作需状态感知能力,而非单次指令跟随。现有机器人系统多依赖昂贵硬件、固定流程、专用设备或面向机器人的接口。本文提出BioProVLA-Agent,一种低成本、协议驱动、视觉增强的多智能体系统,基于视觉-语言-动作(VLA)模型实现生物操作。系统以协议为任务接口,集成协议解析、视觉状态验证与实体执行的闭环流程:定制大语言模型协议代理将协议转为可验证子任务;视觉语言模型-检索增强生成验证代理结合观测、机器人状态、检索知识及成功/失败案例判断就绪与完成状态;VLA实体代理通过轻量策略执行已验证子任务。为应对湿实验中的视觉干扰,开发AugSmolVLA在线增强策略,针对性改善透明器皿、反光、光照变化与过曝问题。在包含15个原子任务、6个复合流程与3个双臂任务的层级基准上评估,涵盖试管装载、分类、废料处理、瓶盖拧动与液体倾倒。在正常与高曝光条件下,AugSmolVLA在精确放置、透明物体操作、复合流程与视觉退化场景中均优于ACT、X-VLA与原始SmolVLA,表明其为实现可访问、协议中心、具备验证能力的生物操作人工智能提供了可行路径。
原文摘要 · Abstract (English)
Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab environments remains challenging. Protocols are often unstructured, labware is frequently transparent or reflective, and multi-step procedures require state-aware execution beyond one-shot instruction following. Existing robotic systems often rely on costly hardware, fixed workflows, dedicated instruments, or robotics-oriented interfaces. Here, we introduce BioProVLA-Agent, an affordable, protocol-driven, vision-enhanced embodied multi-agent system enabled by Vision-Language-Action (VLA) models for biological manipulation. The system uses protocols as the task interface and integrates protocol parsing, visual state verification, and embodied execution in a closed-loop workflow. A Tailored LLM Protocol Agent converts protocols into verifiable subtasks; a VLM-RAG Verification Agent assesses readiness and completion using observations, robot states, retrieved knowledge, and success/failure examples; and a VLA Embodied Agent executes verified subtasks through a lightweight policy. To improve robustness under wet-lab visual perturbations, we develop AugSmolVLA, an online augmentation strategy targeting transparent labware, reflections, illumination shifts, and overexposure. We evaluate the system on a hierarchical benchmark covering 15 atomic tasks, 6 composite workflows, and 3 bimanual tasks, including tube loading, sorting, waste disposal, cap twisting, and liquid pouring. Across normal and high-exposure settings, AugSmolVLA improves execution stability over ACT, X-VLA, and the original SmolVLA, especially for precise placement, transparent-object manipulation, composite workflows, and visually degraded scenes. These results suggest a practical route toward accessible, protocol-centered, and verification-capable embodied AI for biological manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。