arXiv:2501.03374cs.CVcs.AI2025-01被引 6

用扩散模型生成逼真车牌图像,解决数据隐私难题。

License Plate Images Generation with Diffusion Models

  • 基于扩散模型生成车牌图像,模仿真实数据分布。
  • 合成1万张车牌图,使识别准确率提升3%。
  • 适合需要数据增强的智能交通研究者使用。

由于隐私法规(如GDPR)限制,车牌识别(LPR)研究受限于公开数据集的数量。为此,本文提出利用扩散模型生成逼真车牌图像,以缓解数据短缺问题。实验中,模型在乌克兰车牌数据集上训练,并生成1000张合成图像进行人工分类与标注,分析了生成成功率、字符分布及失败模式。研究验证了扩散模型在车牌合成中的有效性,并提供包含10000张图像的公开合成数据集(https://zenodo.org/doi/10.5281/zenodo.13342102)。实验证明,使用伪标签合成数据扩展训练集后,LPR模型准确率较基线提升3%。

原文摘要 · Abstract (English)

Despite the evident practical importance of license plate recognition (LPR), corresponding research is limited by the volume of publicly available datasets due to privacy regulations such as the General Data Protection Regulation (GDPR). To address this challenge, synthetic data generation has emerged as a promising approach. In this paper, we propose to synthesize realistic license plates (LPs) using diffusion models, inspired by recent advances in image and video generation. In our experiments a diffusion model was successfully trained on a Ukrainian LP dataset, and 1000 synthetic images were generated for detailed analysis. Through manual classification and annotation of the generated images, we performed a thorough study of the model output, such as success rate, character distributions, and type of failures. Our contributions include experimental validation of the efficacy of diffusion models for LP synthesis, along with insights into the characteristics of the generated data. Furthermore, we have prepared a synthetic dataset consisting of 10,000 LP images, publicly available at https://zenodo.org/doi/10.5281/zenodo.13342102. Conducted experiments empirically confirm the usefulness of synthetic data for the LPR task. Despite the initial performance gap between the model trained with real and synthetic data, the expansion of the training data set with pseudolabeled synthetic data leads to an improvement in LPR accuracy by 3% compared to baseline.

扩散模型车牌生成数据合成LPR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。