arXiv:2607.23704cs.ROcs.CV2026-07

构建化学自驱动实验室机器人故障分析基准,提升实验可靠性与自主恢复能力。

LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory

论文配图:LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory
图 1 · 摘自论文原文
  • 设计多层级可控故障注入系统,生成超2万条实验轨迹数据
  • 提出六维评估体系,实现故障检测准确率90.83%、定位准确率77.21%
  • 开发专用视觉语言模型,支持实时故障诊断与修复建议

在自驱动实验室中部署具身智能体可加速科学发现,但其可靠性受限于化学实验的不可逆性与安全性。现有研究受制于故障数据稀缺及细粒度评估标准缺失。为此,我们提出LabRobFail,一个以故障为中心的机器人故障分析框架。LabRobFail-Sim在控制、物理和语义层面注入可控故障,构建了包含20,000+轨迹、70+任务场景、5类故障和11种细粒度故障类型的LabRobFail-Data。LabRobFail-Bench评估六项能力:任务理解、故障检测、时间定位、严重性评估、故障分类与可执行修正。我们进一步开发了领域专用的视觉语言模型LabRobFail-VLM,生成结构化故障诊断与恢复指令。在已知环境中,该模型故障检测准确率达90.83%,时间定位准确率为77.21%,显著优于通用视觉语言模型。作为实时监管器集成后,下游任务成功率提升4-16个百分点,验证了细粒度故障理解对闭环恢复与实验室自主性的价值。代码与数据已开源。

原文摘要 · Abstract (English)

The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irreversible and safety-critical nature of chemical experiments. Progress is further hindered by scarce failure data and the lack of fine-grained evaluation protocols. To address these challenges, we introduce LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories. LabRobFail-Sim injects controllable failures at the control, physics, and semantic levels, enabling the construction of LabRobFail-Data, which contains over 20,000 trajectories across 70+ task scenarios, five failure categories, and 11 fine-grained failure types. LabRobFail-Bench evaluates six capabilities spanning task understanding, failure detection, temporal localization, severity assessment, failure classification, and actionable correction. We further develop LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions. On seen environments, it achieves 90.83% failure-detection accuracy and 77.21% temporal-localization accuracy, substantially outperforming general-purpose VLMs. When integrated as a real-time supervisor, it improves downstream task success rates by 4-16 percentage points, demonstrating the value of fine-grained failure understanding for closed-loop recovery and reliable laboratory autonomy. Our code and data are available at https://github.com/Su-ISE-2001/SciRobo

机器人故障分析自驱动实验室视觉语言模型化学自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。