针对匿名化不完整表格数据,提出新转换策略提升机器学习效果。
Learning from Anonymized and Incomplete Tabular Data
- 设计新数据转换方法,保留匿名化语义而非简单丢弃
- 实验证明泛化值比直接删除更有效,提升模型性能
- 适合关注隐私保护下数据利用的开发者与研究者
用户驱动的隐私机制使数据共享具有灵活粒度,导致同一记录中混杂原始、泛化和缺失值。传统机器学习将非原始值视为新类别或缺失,忽略泛化语义。本文提出新型数据转换策略,考虑异构匿名化,并在多个数据集、隐私配置和部署场景下评估其效果。结果表明,泛化值优于纯抑制,最佳预处理策略依赖具体场景,一致的数据表示对下游任务性能至关重要。研究强调,有效学习依赖于对匿名化值的恰当处理。
原文摘要 · Abstract (English)
User-driven privacy allows individuals to control whether and at what granularity their data is shared, leading to datasets that mix original, generalized, and missing values within the same records and attributes. While such representations are intuitive for privacy, they pose challenges for machine learning, which typically treats non-original values as new categories or as missing, thereby discarding generalization semantics. For learning from such tabular data, we propose novel data transformation strategies that account for heterogeneous anonymization and evaluate them alongside standard imputation and LLM-based approaches. We employ multiple datasets, privacy configurations, and deployment scenarios, demonstrating that our method reliably regains utility. Our results show that generalized values are preferable to pure suppression, that the best data preparation strategy depends on the scenario, and that consistent data representations are crucial for maintaining downstream utility. Overall, our findings highlight that effective learning is tied to the appropriate handling of anonymized values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。