解决过程挖掘中局部模型重复多、难筛选的问题。
Framework for Grouping Local Process Models
- 通过上下文与结构相似性对局部模型分组
- 用每组代表模型替代高分模型,减少重复率超80%
- 适合需要清晰理解流程的分析师使用
局部过程模型(LPMs)是流程挖掘中尚未充分探索的概念,可描述事件数据中的顺序、选择、并发和循环模式。近年来,流程挖掘在运营流程分析与优化中表现成功,但忽略完整流程常导致意外发现,使LPMs极具价值。然而,与其它模式挖掘方法类似,LPM发现算法面临模型爆炸与重复问题:可能生成数百甚至上千个模型,其中部分结构或行为高度相似。实践中,分析者难以处理数千个模型,通常仅依赖易获取的高分模型样本。当前观点认为高分模型构成最优样本,但不同应用需求应有不同最优样本。本文表明,若目标是理解流程,以高分模型为样本是劣解,因其重复率过高。为此,提出一种分组框架,通过已知流程模型相似度度量或基于事件日志中数据属性形成的上下文比较,对LPMs进行分组,并为每组选取一个代表性模型构成最优样本。在多个事件日志上的实验显示,相比高分模型样本,分组后样本的重复率降低超过80%,覆盖范围更广。
原文摘要 · Abstract (English)
Local Process Models (LPMs) are an underexplored concept in process mining. LPMs describe patterns in event data considering sequence, choice, concurrency, and loop. In recent years, process mining has proved successful in the analysis and improvement of operational processes. More often than not, surprising findings are found when one does not consider the full process, making LPMs and their discovery highly valuable. However, similar to other pattern mining approaches, LPM discovery algorithms face the problems of model explosion and model repetition, i.e., the algorithms may create hundreds if not thousands of LPMs, and subsets of them are close in structure or behavior. Practically, no analyst would be able to comb through thousands of LPMs leading to using a sample of LPMs that are easily accessible. The current sentiment is that the top-scoring LPMs form the optimal sample to be presented. However, different applications should demand a different optimal sample. With this work, we show that if the goal of the mined LPMs is to understand a process, using the top-scoring LPMs as an optimal sample is a poor choice because of high repetition. We propose a framework for grouping LPMs and creating an optimal sample by taking one representative LPM for each group. We measure similarity between models via established process model similarity measures or by comparing the context in which an LPM appears. The context is formed using data attributes available in the underlying event logs. We demonstrate the usefulness of grouping on multiple event logs by comparing repetition and coverage between samples comprised of the top-scoring models and the representatives of discovered groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。