arXiv:2507.06008cs.CRcs.AI2025-07

通过划分事件数据,提升隐私保护下的流程发现效果。

The Impact of Event Data Partitioning on Privacy-aware Process Discovery

  • 用事件抽象将日志分块,分别匿名化以降低隐私风险。
  • 实验证明,分块后直接后继关系方法的流程发现准确率提升。
  • 适合关注隐私与流程挖掘平衡的系统设计者和研究人员。

信息系统支撑业务流程执行,其事件日志通常包含客户、患者和员工的敏感信息。为应对隐私挑战,可在保留流程发现可用性的前提下对事件日志进行匿名化处理。然而,实用性和隐私性之间的权衡极具挑战:事件日志越复杂,匿名化带来的可用性损失越大。本文提出一种结合匿名化与事件数据分区的流程,利用事件抽象实现日志分块,使每个子日志可独立匿名化。该方法在保障隐私的同时缓解了可用性损失。我们通过三个真实世界事件日志和两种流程发现技术,评估了事件分区对两种匿名化技术的影响。结果表明,事件分区能显著提升基于直接后继关系的匿名化方法在流程发现中的实用性。

原文摘要 · Abstract (English)

Information systems support the execution of business processes. The event logs of these executions generally contain sensitive information about customers, patients, and employees. The corresponding privacy challenges can be addressed by anonymizing the event logs while still retaining utility for process discovery. However, trading off utility and privacy is difficult: the higher the complexity of event log, the higher the loss of utility by anonymization. In this work, we propose a pipeline that combines anonymization and event data partitioning, where event abstraction is utilized for partitioning. By leveraging event abstraction, event logs can be segmented into multiple parts, allowing each sub-log to be anonymized separately. This pipeline preserves privacy while mitigating the loss of utility. To validate our approach, we study the impact of event partitioning on two anonymization techniques using three real-world event logs and two process discovery techniques. Our results demonstrate that event partitioning can bring improvements in process discovery utility for directly-follows-based anonymization techniques.

隐私保护流程发现事件日志

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。