arXiv:2601.04641cs.CRcs.CL2026-01

用自适应差分隐私保护用户数据,同时提升机器文本检测准确率。

DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization

  • 通过两阶段噪声估计动态分配隐私预算,分别处理数值与文本实体。
  • 在MGTBench-2.0上实现接近完美的检测精度,优于非私有基线。
  • 揭示了差分隐私噪声反而增强人机文本区分能力的反直觉现象。

机器生成文本(MGT)检测系统需处理敏感用户数据,导致作者身份验证与隐私保护之间的根本矛盾。标准匿名化方法常破坏语言流畅性,而严格的差分隐私(DP)机制则通常削弱检测所需统计信号。为此,我们提出DP-MGTD框架,集成自适应差分隐私实体净化算法。该方法采用两阶段机制,对频率进行噪声估计并动态校准隐私预算,分别使用拉普拉斯和指数机制处理数值与文本实体。关键发现:应用DP噪声会通过暴露对扰动的不同敏感模式,反而增强人机文本的可区分性。在MGTBench-2.0数据集上的大量实验表明,本方法在满足严格隐私保障的同时,实现了近乎完美的检测准确率,显著优于非私有基线。

原文摘要 · Abstract (English)

The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation. Standard anonymization techniques often disrupt linguistic fluency, while rigorous Differential Privacy (DP) mechanisms typically degrade the statistical signals required for accurate detection. To resolve this dilemma, we propose \textbf{DP-MGTD}, a framework incorporating an Adaptive Differentially Private Entity Sanitization algorithm. Our approach utilizes a two-stage mechanism that performs noisy frequency estimation and dynamically calibrates privacy budgets, applying Laplace and Exponential mechanisms to numerical and textual entities respectively. Crucially, we identify a counter-intuitive phenomenon where the application of DP noise amplifies the distinguishability between human and machine text by exposing distinct sensitivity patterns to perturbation. Extensive experiments on the MGTBench-2.0 dataset show that our method achieves near-perfect detection accuracy, significantly outperforming non-private baselines while satisfying strict privacy guarantees.

隐私保护文本检测差分隐私机器生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。