arXiv:2509.12375cs.LG2025-09中稿 · ITSC 2025

用扩散模型生成真实驾驶数据并修复异常,提升小样本数据质量。

Diffusion-Based Generation and Imputation of Driving Scenarios from Limited Vehicle CAN Data

  • 结合自回归与非自回归方法,改进扩散模型生成时序数据。
  • 生成数据物理正确性超越原始训练数据,驾驶行为合理。
  • 可有效修复真实数据中不合理的驾驶片段,适合自动驾驶数据增强。

在仅有少量且含噪声的时间序列数据上训练深度学习模型极具挑战。扩散模型在生成逼真合成数据和修复缺失/异常数据方面表现优异。本文聚焦于生成真实可靠的汽车时序数据,将去噪扩散概率模型(DDPM)应用于具有长期依赖且样本稀少的车辆CAN数据集。提出一种混合生成方法,融合自回归与非自归模型优势,并针对两种近期提出的时序生成DDPM架构进行多项改进。设计三种评估指标,量化生成数据的物理合理性与轨迹符合度。最佳模型在物理正确性上优于原始训练数据,且表现出自然驾驶行为。进一步利用该模型成功修复训练数据中物理不合理区域,显著提升数据质量。

原文摘要 · Abstract (English)

Training deep learning methods on small time series datasets that also include corrupted samples is challenging. Diffusion models have shown to be effective to generate realistic and synthetic data, and correct corrupted samples through imputation. In this context, this paper focuses on generating synthetic yet realistic samples of automotive time series data. We show that denoising diffusion probabilistic models (DDPMs) can effectively solve this task by applying them to a challenging vehicle CAN-dataset with long-term data and a limited number of samples. Therefore, we propose a hybrid generative approach that combines autoregressive and non-autoregressive techniques. We evaluate our approach with two recently proposed DDPM architectures for time series generation, for which we propose several improvements. To evaluate the generated samples, we propose three metrics that quantify physical correctness and test track adherence. Our best model is able to outperform even the training data in terms of physical correctness, while showing plausible driving behavior. Finally, we use our best model to successfully impute physically implausible regions in the training data, thereby improving the data quality.

时序生成扩散模型自动驾驶数据修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。