arXiv:2509.12950cs.IRcs.DS2025-09中稿 · NetMob 2025 Data C…被引 1

提出新方法保护真实人口隐私,而非仅匿名参与者。

Protecting participants or population? Comparison of k-anonymous Origin-Destination matrices

  • 设计新算法ODkAnon,在保证效率的同时实现人口级k匿名
  • 实验证明新方法在保持数据可用性上优于传统地理泛化法
  • 适合关注真实人群隐私的交通、城市规划研究者

出行起讫(OD)矩阵是研究用户移动行为的核心工具,用于总结个体在地理区域间的流动情况。为兼顾代表性与隐私风险,需将区域划分足够细致。本研究基于NetMob2025挑战赛数据集,该数据具有丰富社会人口统计信息,可按人群分段生成多类OD矩阵。更重要的是,每位参与者不仅是数据记录,更是对真实人群的统计加权代理。这一特性推动了匿名化范式的根本转变——从保护个体参与者转向保护真实人群的推断身份。本文旨在构建并比较针对调查参与者和整个真实人口均满足k匿名性的OD矩阵。我们对比了多种传统匿名化方法,包括基于层级的区域泛化(ATG、OIGH)及经典Mondrian方法,并引入一种新算法ODkAnon——一种贪心策略,兼顾速度与质量。与以往仅关注数据集隐私的方法不同,本工作致力于生成兼具社会人口分段特征、且对真实人群实现k匿名的隐私保护型OD矩阵。

原文摘要 · Abstract (English)

Origin-Destination (OD) matrices are a core component of research on users' mobility and summarize how individuals move between geographical regions. These regions should be small enough to be representative of user mobility, without incurring substantial privacy risks. There are two added values of the NetMob2025 challenge dataset. Firstly, the data is extensive and contains a lot of socio-demographic information that can be used to create multiple OD matrices, based on the segments of the population. Secondly, a participant is not merely a record in the data, but a statistically weighted proxy for a segment of the real population. This opens the door to a fundamental shift in the anonymization paradigm. A population-based view of privacy is central to our contribution. By adjusting our anonymization framework to account for representativeness, we are also protecting the inferred identity of the actual population, rather than survey participants alone. The challenge addressed in this work is to produce and compare OD matrices that are k-anonymous for survey participants and for the whole population. We compare several traditional methods of anonymization to k-anonymity by generalizing geographical areas. These include generalization over a hierarchy (ATG and OIGH) and the classical Mondrian. To this established toolkit, we add a novel method, i.e., ODkAnon, a greedy algorithm aiming at balancing speed and quality. Unlike previous approaches, which primarily address the privacy aspects of the given datasets, we aim to contribute to the generation of privacy-preserving OD matrices enriched with socio-demographic segmentation that achieves k-anonymity on the actual population.

隐私保护出行矩阵人口匿名

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。