arXiv:2509.08820cs.RO2025-09中稿 · CoRL被引 12

让机器人安全完成复杂化学实验,成功率提升23.57%。

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation

  • 用视觉语言模型分解任务并生成视觉提示,指导机械臂操作。
  • 相比顶尖系统,成功率提高23.57%,合规率提升0.298。
  • 适合需要长时序、高安全性的自动化化学实验场景。

机器人化学家有望解放科研人员并加速科学发现,但当前仍处于早期阶段。化学实验涉及长期任务和危险、易变形物质,成功不仅需完成操作,还需严格遵守实验规范。为此,我们提出RoboChemist,一种融合视觉-语言模型(VLM)与视觉-语言-动作模型(VLA)的双环框架。不同于依赖深度感知且难以处理透明器皿的VLM系统,以及缺乏语义反馈的VLA系统,本方法利用VLM实现:(1) 任务规划,将复杂任务分解为基本动作;(2) 视觉提示生成,引导VLA模型;(3) 任务执行与合规性监控。特别地,我们设计了接受图像目标输入的VLA接口,实现精准的目标条件控制。系统成功执行基础操作与多步化学协议,结果表明,平均成功率比最先进VLA基线高出23.57%,合规率平均提升0.298,并展现出对新物体和任务的强大泛化能力。

原文摘要 · Abstract (English)

Robotic chemists promise to both liberate human experts from repetitive tasks and accelerate scientific discovery, yet remain in their infancy. Chemical experiments involve long-horizon procedures over hazardous and deformable substances, where success requires not only task completion but also strict compliance with experimental norms. To address these challenges, we propose \textit{RoboChemist}, a dual-loop framework that integrates Vision-Language Models (VLMs) with Vision-Language-Action (VLA) models. Unlike prior VLM-based systems (e.g., VoxPoser, ReKep) that rely on depth perception and struggle with transparent labware, and existing VLA systems (e.g., RDT, pi0) that lack semantic-level feedback for complex tasks, our method leverages a VLM to serve as (1) a planner to decompose tasks into primitive actions, (2) a visual prompt generator to guide VLA models, and (3) a monitor to assess task success and regulatory compliance. Notably, we introduce a VLA interface that accepts image-based visual targets from the VLM, enabling precise, goal-conditioned control. Our system successfully executes both primitive actions and complete multi-step chemistry protocols. Results show 23.57% higher average success rate and a 0.298 average increase in compliance rate over state-of-the-art VLA baselines, while also demonstrating strong generalization to objects and tasks.

机器人化学视觉语言自动化实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。