arXiv:2602.11015cs.CRcs.AI2026-02

提出几何框架CVPL,量化数据脱敏后的真实链接风险。

CVPL: A Geometric Framework for Post-Hoc Linkage Risk Assessment in Protected Tabular Data

  • 将链接分析建模为阻断、向量化、投影、相似性评估的流程
  • 实测显示k匿名合规数据仍存在显著链接风险,部分源于行为模式
  • 可解释特征贡献,适合隐私评估与脱敏方案对比

正式隐私度量提供合规保障但难以量化真实链接可能性。我们提出CVPL(聚类-向量-投影链接)几何框架,用于对原始与脱敏表格数据间的链接风险进行事后评估。CVPL将链接分析表示为阻断、向量化、潜在投影和相似性评估的算子流水线,生成连续的、场景依赖的风险估计,而非二元合规判断。在显式威胁模型下形式化定义CVPL,引入阈值感知风险曲面R(λ, τ),捕捉保护强度与攻击者严格程度的联合影响。提出具有单调性保证的渐进阻断策略,实现任意时间的风险下界估计。证明经典Fellegi-Sunter链接是CVPL在严格假设下的特例,且这些假设的违反会导致系统性过链接偏差。在19种保护配置下对10,000条记录的实证验证表明,形式上的k匿名合规可能与显著的实证链接性并存,其中大量风险源于非准标识符的行为模式。CVPL提供可解释诊断,识别推动链接可行性的特征,支持隐私影响评估、保护机制比较及效用-风险权衡分析。

原文摘要 · Abstract (English)

Formal privacy metrics provide compliance-oriented guarantees but often fail to quantify actual linkability in released datasets. We introduce CVPL (Cluster-Vector-Projection Linkage), a geometric framework for post-hoc assessment of linkage risk between original and protected tabular data. CVPL represents linkage analysis as an operator pipeline comprising blocking, vectorization, latent projection, and similarity evaluation, yielding continuous, scenario-dependent risk estimates rather than binary compliance verdicts. We formally define CVPL under an explicit threat model and introduce threshold-aware risk surfaces, R(lambda, tau), that capture the joint effects of protection strength and attacker strictness. We establish a progressive blocking strategy with monotonicity guarantees, enabling anytime risk estimation with valid lower bounds. We demonstrate that the classical Fellegi-Sunter linkage emerges as a special case of CVPL under restrictive assumptions, and that violations of these assumptions can lead to systematic over-linking bias. Empirical validation on 10,000 records across 19 protection configurations demonstrates that formal k-anonymity compliance may coexist with substantial empirical linkability, with a significant portion arising from non-quasi-identifier behavioral patterns. CVPL provides interpretable diagnostics identifying which features drive linkage feasibility, supporting privacy impact assessment, protection mechanism comparison, and utility-risk trade-off analysis.

隐私评估数据脱敏链接风险可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。