arXiv:2604.16345cs.HCcs.AI2026-04

用AI把实验室口头经验转成可执行的指导,让新手也能安全操作。

Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support

  • 结合第一视角视频与多模态AI,从实验录像中提取未写入手册的操作细节
  • 通过RAG和双层安全设计,确保回答基于真实手册且不胡编乱造
  • 专家评估显示生成建议既实用又安全,适合教学与探索性实验室使用

尽管材料信息学推动了自驱动实验室的发展,但许多教育及探索性实验室仍依赖人工实验。在特定实验环境中,仅靠正式文档往往不足以保障安全可靠的操作。我们称这种正式文档与实际操作之间的差距为‘实验最后一公里’,其核心是难以书面化的现场经验,包括本地规则、日常检查、操作细节和安全行为等。本概念验证研究开发了一种人机协作的AI助手,融合第一视角实验视频、多模态AI与检索增强生成(RAG)。以粉末X射线衍射实验和学生录制的视频数据为输入,系统从记录的操作中提取物理技巧与可听确认信号,这些常被传统手册忽略。随后,基于提取的知识生成有依据的响应。为降低无支持输出风险,系统采用双层安全设计:通过RAG限制信息源,并施加严格系统提示约束。导师评估显示,对手册覆盖问题的回答与预期一致;对超出范围的问题,系统能正确拒绝回答,表明幻觉风险降低。专家评估进一步证实,生成的咨询报告具有实用性(3.25/4.00)与安全性(4.00/4.00)。结果表明,该框架在明确人类监督下,可实现对实验最后一公里的弥合,使AI辅助实验室实践而非替代人类判断。

原文摘要 · Abstract (English)

While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain standard in many educational and exploratory research laboratories. In specific lab settings, formal documentation alone is often insufficient for safe and reliable operation. We refer to the gap between formal documentation and reliable execution in such settings as the experimental last mile; this gap mainly involves site-specific operational know-how, including local rules, routine checks, procedural details, and safety-conscious actions that are can be verbalizable but are often under-documented in standard manuals. In this proof-of-concept study, we developed a human-in-the-loop AI assistant that combines first-person experimental video, multimodal AI, and retrieval-augmented generation (RAG). Using powder X-ray diffraction experiments and student-recorded video data as inputs, the system extracts site-specific laboratory knowledge from recorded procedures, including physical techniques and audible confirmation that conventional manuals could omit. It then provides grounded responses based on the resulting manual. To reduce the risk of unsupported outputs, the system employs a two-layer safety design: source restriction through RAG and strict system-prompt constraints. Instructor-based evaluation showed alignment with expected guidance for questions covered by the manual. For out-of-scope queries, the system appropriately refused to answer, indicating a reduced risk of hallucination. Expert evaluation further indicated that the generated advisory reports were useful and safe (utility: 3.25/4.00; safety: 4.00/4.00). These results suggest the feasibility of a framework for bridging the experimental last mile in which AI supports laboratory practice under explicit human supervision rather than replacing human judgment.

实验室自动化AI助手多模态安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。