arXiv:2601.22754cs.CVcs.AI2026-01中稿 · the 23rd IFAC Worl…被引 1

用视觉语言模型自动解析工业故障排查图,提升维修支持效率

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

  • 结合图文信息的视觉语言模型解析流程图式故障指南
  • 增强提示策略使模型更敏感于布局模式,提升结构提取准确率
  • 为工厂现场系统部署提供实用选型参考

工业故障排查指南以类似流程图的图表形式编码诊断步骤,其空间布局与技术语言共同传达含义。为将此知识整合至辅助车间人员诊断和修复设备问题的操作支持系统中,需先将其提取并结构化以便机器理解。然而,人工提取耗时且易出错。视觉语言模型(VLM)有望通过联合解析视觉与文本信息来自动化该过程,但其在处理此类指南上的表现尚不明确。本文评估了两种VLM在结构化知识提取任务中的表现,比较了标准指令引导与一种提示增强方法——后者引入故障排查布局模式作为上下文线索。结果揭示不同模型在布局敏感性与语义鲁棒性间的权衡,为实际部署提供了依据。

原文摘要 · Abstract (English)

Industrial troubleshooting guides encode diagnostic procedures in flowchart-like diagrams where spatial layout and technical language jointly convey meaning. To integrate this knowledge into operator support systems, which assist shop-floor personnel in diagnosing and resolving equipment issues, the information must first be extracted and structured for machine interpretation. However, when performed manually, this extraction is labor-intensive and error-prone. Vision Language Models offer potential to automate this process by jointly interpreting visual and textual meaning, yet their performance on such guides remains underexplored. This paper evaluates two VLMs on extracting structured knowledge, comparing two prompting strategies: standard instruction-guided versus an augmented approach that cues troubleshooting layout patterns. Results reveal model-specific trade-offs between layout sensitivity and semantic robustness, informing practical deployment decisions.

视觉语言模型知识提取工业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。