arXiv:2507.18293cs.LG2025-07

用统计增强与孪生学习提升流程预测的泛化能力

Leveraging Data Augmentation and Siamese Learning for Predictive Process Monitoring

  • 基于控制流语义生成真实流程变体,增强数据多样性
  • 在真实日志上实现优于或媲美当前最优的预测性能
  • 适合数据少、标注难的工业流程预测场景

预测性流程监控(PPM)可根据事件日志预测正在进行的业务流程实例的未来事件或结果。然而,深度学习方法常受限于真实世界事件日志的低多样性与小规模。为此,我们提出SiamSA-PPM,一种结合孪生学习与统计增强的自监督学习框架。该方法设计三种基于统计的变换策略,利用控制流语义和常见行为模式生成语义合理的新轨迹变体。这些增强视图被用于孪生网络中,以无监督方式学习流程前缀的可泛化表示。在真实事件日志上的大量实验表明,SiamSA-PPM在下一活动和最终结果预测任务中均达到竞争性或更优性能。结果还显示,统计增强显著优于随机变换,有效提升数据变异性,验证了SiamSA-PPM在流程预测数据增广中的潜力。

原文摘要 · Abstract (English)

Predictive Process Monitoring (PPM) enables forecasting future events or outcomes of ongoing business process instances based on event logs. However, deep learning PPM approaches are often limited by the low variability and small size of real-world event logs. To address this, we introduce SiamSA-PPM, a novel self-supervised learning framework that combines Siamese learning with Statistical Augmentation for Predictive Process Monitoring. It employs three novel statistically grounded transformation methods that leverage control-flow semantics and frequent behavioral patterns to generate realistic, semantically valid new trace variants. These augmented views are used within a Siamese learning setup to learn generalizable representations of process prefixes without the need for labeled supervision. Extensive experiments on real-life event logs demonstrate that SiamSA-PPM achieves competitive or superior performance compared to the SOTA in both next activity and final outcome prediction tasks. Our results further show that statistical augmentation significantly outperforms random transformations and improves variability in the data, highlighting SiamSA-PPM as a promising direction for training data enrichment in process prediction.

流程预测自监督学习数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。