arXiv:2409.03669stat.MLcs.AI2024-09被引 3

提出可控生成方法,用于评估高维工艺漂移检测算法性能。

A method to benchmark high-dimensional process drift detection

  • 构建理论框架,可控生成多变量工艺曲线数据
  • 引入时间曲线下面积评分,量化模型识别漂移段能力
  • 发现现有算法在多漂移段场景下普遍表现不佳

工艺曲线是来自制造过程的多变量有限时间序列数据。本文研究用于检测工艺曲线数据集漂移的机器学习方法。提出一种理论框架,可受控地合成工艺曲线,以基准化机器学习算法在工艺漂移检测中的表现。引入一种名为时间曲线下面积(temporal area under the curve)的评估指标,用于量化机器学习模型揭示漂移段曲线的能力。最后,基于该框架生成的合成数据,对主流机器学习方法进行了基准比较,结果表明现有算法在包含多个漂移段的数据集上常表现不佳。

原文摘要 · Abstract (English)

Process curves are multivariate finite time series data coming from manufacturing processes. This paper studies machine learning that detect drifts in process curve datasets. A theoretic framework to synthetically generate process curves in a controlled way is introduced in order to benchmark machine learning algorithms for process drift detection. An evaluation score, called the temporal area under the curve, is introduced, which allows to quantify how well machine learning models unveil curves belonging to drift segments. Finally, a benchmark study comparing popular machine learning approaches on synthetic data generated with the introduced framework is presented that shows that existing algorithms often struggle with datasets containing multiple drift segments.

漂移检测制造优化算法评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。