用合成数据保护智能汽车隐私,有效防追踪还保数据可用。
Driving Privacy Forward: Mitigating Information Leakage within Smart Vehicles through Synthetic Data Generation
- 用变分自编码器生成百万级车辆传感器合成数据
- 90.1%统计相似度,78%分类准确率,同时阻止驾驶者画像
- 为车联网隐私防护提供可复现的技术方案,适合安全研究者
智能汽车产生大量数据,其中许多敏感信息易遭隐私泄露。随着攻击者越来越多地利用匿名化元数据进行驾驶者画像,亟需在不阻碍创新的前提下解决信息泄漏问题。合成数据成为缓解隐私风险的有力工具,可在保留真实数据关系的同时降低敏感信息暴露风险。本文提出一个涵盖14类车内传感器的综合分类体系,识别潜在攻击并评估其脆弱性。针对最脆弱信号,基于Passive Vehicular Sensor(PVS)数据集,使用表格式变分自编码器(TVAE)生成超百万条合成数据。通过三个核心指标——保真度、实用性与隐私性进行评估,结果表明:合成数据在原任务上保持90.1%的统计相似度和78%的分类准确率,同时成功防止驾驶者身份推理。代码已开源。
原文摘要 · Abstract (English)
Smart vehicles produce large amounts of data, much of which is sensitive and at risk of privacy breaches. As attackers increasingly exploit anonymised metadata within these datasets to profile drivers, it's important to find solutions that mitigate this information leakage without hindering innovation and ongoing research. Synthetic data has emerged as a promising tool to address these privacy concerns, as it allows for the replication of real-world data relationships while minimising the risk of revealing sensitive information. In this paper, we examine the use of synthetic data to tackle these challenges. We start by proposing a comprehensive taxonomy of 14 in-vehicle sensors, identifying potential attacks and categorising their vulnerability. We then focus on the most vulnerable signals, using the Passive Vehicular Sensor (PVS) dataset to generate synthetic data with a Tabular Variational Autoencoder (TVAE) model, which included over 1 million data points. Finally, we evaluate this against 3 core metrics: fidelity, utility, and privacy. Our results show that we achieved 90.1% statistical similarity and 78% classification accuracy when tested on its original intent while also preventing the profiling of the driver. The code can be found at https://github.com/krish-parikh/Synthetic-Data-Generation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。