arXiv:2411.10172cs.CLcs.AI2024-11被引 3

用自动化方法从半导体文档中提取因果关系,提升故障预测与流程优化能力

Increasing the Accessibility of Causal Domain Knowledge via Causal Information Extraction Methods: A Case Study in the Semiconductor Manufacturing Industry

  • 提出单阶段和多阶段序列标注方法,从工业文本中抽取因果信息
  • 在FMEA文档上达93%准确率,在演示文稿上达73%准确率
  • 强调领域适配语言模型与微调对效果的关键作用,适合工业知识工程

从文本数据中提取因果信息对工业领域至关重要,有助于识别和缓解潜在故障、提升工艺效率、推动质量改进并应对各类运营挑战。本文针对半导体制造行业的实际文档,研究了自动化因果信息抽取方法。提出了两种方法:单阶段序列标注(SST)和多阶段序列标注(MST),并在某半导体公司的真实文档(包括演示文稿和FMEA文档)上进行了评估。研究还探讨了表示学习对下游任务的影响。案例研究表明,所提出的MST方法在非结构化程度较高的FMEA文档上表现优异,达到93% F1分数;在演示文稿文本上也实现了73% F1分数。此外,研究强调选择更契合领域的语言模型及进行领域内微调的重要性。

原文摘要 · Abstract (English)

The extraction of causal information from textual data is crucial in the industry for identifying and mitigating potential failures, enhancing process efficiency, prompting quality improvements, and addressing various operational challenges. This paper presents a study on the development of automated methods for causal information extraction from actual industrial documents in the semiconductor manufacturing industry. The study proposes two types of causal information extraction methods, single-stage sequence tagging (SST) and multi-stage sequence tagging (MST), and evaluates their performance using existing documents from a semiconductor manufacturing company, including presentation slides and FMEA (Failure Mode and Effects Analysis) documents. The study also investigates the effect of representation learning on downstream tasks. The presented case study showcases that the proposed MST methods for extracting causal information from industrial documents are suitable for practical applications, especially for semi structured documents such as FMEAs, with a 93\% F1 score. Additionally, MST achieves a 73\% F1 score on texts extracted from presentation slides. Finally, the study highlights the importance of choosing a language model that is more aligned with the domain and in-domain fine-tuning.

因果抽取工业AI半导体文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。