用概率模型提升流程日志聚类,让每簇更清晰易懂。
Model-driven Stochastic Trace Clustering
- 基于活动转移概率的熵相关性度量,指导日志轨迹归类。
- 聚类后模型在概率一致性上优于传统方法,结构更简洁。
- 适合关注真实执行概率动态的流程分析场景。
过程发现算法可自动从事件日志中提取过程模型,但高变异性常导致模型复杂难懂。为缓解此问题,轨迹聚类技术将过程执行分组,每组由一个更简单、更易理解的过程模型表示。模型驱动的聚类通过轨迹与簇内特定过程模型的符合程度进行分配。然而,现有方法多依赖无模型或非概率模型,忽略活动与转换的频率或概率,限制了对真实执行动态的捕捉能力。本文提出一种新型模型驱动的轨迹聚类方法,优化每个簇内的随机过程模型。该方法采用基于直接跟随概率的熵相关性(entropic relevance)作为随机符合度量,引导轨迹分配。使聚类决策同时考虑结构对齐性与轨迹源自特定随机过程模型的可能性。方法计算高效,随输入线性增长,显著提升模型可解释性,生成具有更清晰控制流模式的簇。在多个公开真实数据集上的实验表明,本方法在随机一致性与图简洁性上表现更优,而传统适应度指标显示存在权衡,凸显其在随机过程分析中的独特价值。
原文摘要 · Abstract (English)
Process discovery algorithms automatically extract process models from event logs, but high variability often results in complex and hard-to-understand models. To mitigate this issue, trace clustering techniques group process executions into clusters, each represented by a simpler and more understandable process model. Model-driven trace clustering improves on this by assigning traces to clusters based on their conformity to cluster-specific process models. However, most existing clustering techniques rely on either no process model discovery, or non-stochastic models, neglecting the frequency or probability of activities and transitions, thereby limiting their capability to capture real-world execution dynamics. We propose a novel model-driven trace clustering method that optimizes stochastic process models within each cluster. Our approach uses entropic relevance, a stochastic conformance metric based on directly-follows probabilities, to guide trace assignment. This allows clustering decisions to consider both structural alignment with a cluster's process model and the likelihood that a trace originates from a given stochastic process model. The method is computationally efficient, scales linearly with input size, and improves model interpretability by producing clusters with clearer control-flow patterns. Extensive experiments on public real-life datasets demonstrate that while our method yields superior stochastic coherence and graph simplicity, traditional fitness metrics reveal a trade-off, highlighting the specific utility of our approach for stochastic process analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。