用生成模型释放隐私保护的事件日志,兼顾高隐私与高可用性。
Releasing Differentially Private Event Logs Using Generative Models
- 基于GAN和扩散模型生成私密化轨迹变体,隐式学习分布。
- 在罕见轨迹占比高的场景下仍保持强隐私保护与数据可用性。
- 适合需要大规模部署且对隐私要求严苛的工业级流程挖掘应用。
近年来,流程挖掘与自动化事件数据分析在工业界广泛应用,但事件数据中包含敏感信息带来的隐私担忧日益突出。现有研究主要通过差分隐私为流程挖掘核心使用的轨迹变体提供可量化的隐私保障,但在面对高频罕见变体时仍显不足。本文提出两种基于训练生成模型的新方法:一是使用生成对抗网络(GANs)从隐私化隐式变体分布中采样(TraVaG);二是利用去噪扩散概率模型(Denoising Diffusion Probabilistic Models)通过训练马尔可夫链从噪声重构人工轨迹变体。两种方法均支持工业规模应用,显著提升隐私保障,尤其在罕见变体较多的场景中表现优异。它们克服了传统方法如限制轨迹长度或引入虚假变体的缺陷。在真实事件数据上的实验表明,本方法在隐私保护与数据效用之间达到更优平衡,优于现有最先进技术。
原文摘要 · Abstract (English)
In recent years, the industry has been witnessing an extended usage of process mining and automated event data analysis. Consequently, there is a rising significance in addressing privacy apprehensions related to the inclusion of sensitive and private information within event data utilized by process mining algorithms. State-of-the-art research mainly focuses on providing quantifiable privacy guarantees, e.g., via differential privacy, for trace variants that are used by the main process mining techniques, e.g., process discovery. However, privacy preservation techniques designed for the release of trace variants are still insufficient to meet all the demands of industry-scale utilization. Moreover, ensuring privacy guarantees in situations characterized by a high occurrence of infrequent trace variants remains a challenging endeavor. In this paper, we introduce two novel approaches for releasing differentially private trace variants based on trained generative models. With TraVaG, we leverage \textit{Generative Adversarial Networks} (GANs) to sample from a privatized implicit variant distribution. Our second method employs \textit{Denoising Diffusion Probabilistic Models} that reconstruct artificial trace variants from noise via trained Markov chains. Both methods offer industry-scale benefits and elevate the degree of privacy assurances, particularly in scenarios featuring a substantial prevalence of infrequent variants. Also, they overcome the shortcomings of conventional privacy preservation techniques, such as bounding the length of variants and introducing fake variants. Experimental results on real-life event data demonstrate that our approaches surpass state-of-the-art techniques in terms of privacy guarantees and utility preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。