arXiv:2502.08856cs.LG2025-02KDD被引 2

评估生成模型在交通数据上的表现,发现现有方法结构保真度不足。

A Systematic Evaluation of Generative Models on Tabular Transportation Data

  • 针对交通数据网络结构设计新图相似性度量
  • 实测多种生成模型在真实纽约出租车数据上表现参差不齐
  • 适合关注交通数据隐私与结构还原的研究者

共享大规模交通数据有助于交通规划与政策制定,但同时也带来安全与隐私风险,因数据可能包含个人位置等敏感信息。基于真实交通数据的合成数据生成为解决该问题提供了潜在方案,可在保护隐私的同时保留数据效用。尽管已有多种合成数据生成技术,但大多未针对交通数据特有的网络结构(如行程构成的路网)进行优化。本文以纽约市出租车数据为例,系统评估了主流表格型生成模型的表现。除了传统的分布相似性、覆盖度和隐私保护指标外,提出一种专为交通数据设计的图结构度量,用于评估真实与合成交通网络间的结构功能对齐程度。同时改进了隐私度量方法,弥补常用指标的缺陷。实验结果表明,现有生成模型在实际交通数据应用中表现远不如文献宣称一致;且新图度量揭示合成数据与真实数据间存在显著差距。研究强调需开发专门适配交通等新兴领域特性的生成模型。

原文摘要 · Abstract (English)

The sharing of large-scale transportation data is beneficial for transportation planning and policymaking. However, it also raises significant security and privacy concerns, as the data may include identifiable personal information, such as individuals' home locations. To address these concerns, synthetic data generation based on real transportation data offers a promising solution that allows privacy protection while potentially preserving data utility. Although there are various synthetic data generation techniques, they are often not tailored to the unique characteristics of transportation data, such as the inherent structure of transportation networks formed by all trips in the datasets. In this paper, we use New York City taxi data as a case study to conduct a systematic evaluation of the performance of widely used tabular data generative models. In addition to traditional metrics such as distribution similarity, coverage, and privacy preservation, we propose a novel graph-based metric tailored specifically for transportation data. This metric evaluates the similarity between real and synthetic transportation networks, providing potentially deeper insights into their structural and functional alignment. We also introduced an improved privacy metric to address the limitations of the commonly-used one. Our experimental results reveal that existing tabular data generative models often fail to perform as consistently as claimed in the literature, particularly when applied to transportation data use cases. Furthermore, our novel graph metric reveals a significant gap between synthetic and real data. This work underscores the potential need to develop generative models specifically tailored to take advantage of the unique characteristics of emerging domains, such as transportation.

生成模型交通数据隐私保护图结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。