arXiv:2511.11847cs.IRcs.AI2025-11被引 2

打造工业安全多模态聊天机器人,精准高效指导操作员。

A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches

  • 用大模型结合检索增强生成,从专业文档中获取安全知识。
  • 最高准确率达86.66%,单次查询成本仅0.005美元,延迟10.04秒。
  • 提供可复用的评测基准与系统化设计方法,适合制造业安全培训。

现代制造环境中的工人安全仍是重大挑战。产业5.0倡导以人为本的生产模式。基于设计科学方法论,我们识别出下一代安全培训系统的三大核心需求:高精度、低延迟和低成本。本文提出一种基于大语言模型的多模态聊天机器人,通过检索增强生成(RAG)技术,使回答基于精心整理的法规与技术文档。为评估方案,我们构建了一个领域专用的基准,包含针对三类典型设备(桥式立铣床、哈斯TL-1数控车床、优傲UR5e协作机器人)的专家验证问答对。采用全因子实验设计,测试了24种RAG配置,并以正确性、延迟和成本进行自动化评估。表现最优的两个配置由十位行业专家与学术研究人员进行人工评测。结果表明,检索策略与模型配置显著影响性能。最佳配置部署后达到86.66%准确率,平均查询成本0.005美元,端到端延迟10.04秒(从提交问题到完整指令交付)。该延迟在实际应用中可接受。本研究贡献包括:一个开源、领域相关的安全培训聊天机器人;一个经验证的AI辅助安全指导评测基准;以及一套面向产业5.0环境的AI教学系统设计与评估方法。

原文摘要 · Abstract (English)

Ensuring worker safety remains a critical challenge in modern manufacturing environments. Industry 5.0 reorients the prevailing manufacturing paradigm toward more human-centric operations. Using a design science research methodology, we identify three essential requirements for next-generation safety training systems: high accuracy, low latency, and low cost. We introduce a multimodal chatbot powered by large language models that meets these design requirements. The chatbot uses retrieval-augmented generation to ground its responses in curated regulatory and technical documentation. To evaluate our solution, we developed a domain-specific benchmark of expert-validated question and answer pairs for three representative machines: a Bridgeport manual mill, a Haas TL-1 CNC lathe, and a Universal Robots UR5e collaborative robot. We tested 24 RAG configurations using a full-factorial design and assessed them with automated evaluations of correctness, latency, and cost. Our top 2 configurations were then evaluated by ten industry experts and academic researchers. Our results show that retrieval strategy and model configuration have a significant impact on performance. The top configuration, selected for chatbot deployment, achieved an accuracy of 86.66%, an average cost of $0.005 per query, and an average end-to-end latency of 10.04 seconds. This latency is practical for delivering a complete safety instruction and is measured from query submission to full instruction delivery rather than generation onset. Overall, our work provides three contributions: an open-source, domain-grounded safety training chatbot; a validated benchmark for evaluating AI-assisted safety instruction; and a systematic methodology for designing and assessing AI-enabled instructional and immersive safety training systems for Industry 5.0 environments.

工业安全多模态RAGLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。