arXiv:2508.10397cs.CVcs.AI2025-08

用姿态驱动的增强方法,在数据少时提升驾驶分心检测效果。

PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

  • 根据驾驶者姿态生成多样样本,再用视觉语言模型筛选高质量数据。
  • 在少样本场景下显著提升模型泛化能力,准确率明显优于基线。
  • 适合数据稀缺的自动驾驶安全系统开发,尤其关注真实场景适应性。

驾驶分心检测对提升交通安全性、减少交通事故至关重要。然而,现有模型在实际部署中常因泛化能力下降而表现不佳,主要源于真实环境中数据标注成本高带来的少样本学习挑战,以及训练数据与实际部署条件之间的显著领域偏移。为此,本文提出一种姿态驱动的质量可控数据增强框架(PQ-DAF),利用视觉语言模型进行样本筛选,以低成本扩充训练数据并增强跨域鲁棒性。具体而言,采用渐进式条件扩散模型(PCDMs)精准捕捉关键驾驶者姿态特征,并生成多样化训练样本;随后构建基于CogVLM视觉语言模型的样本质量评估模块,依据置信度阈值过滤低质量合成样本,确保增强数据集的可靠性。大量实验表明,PQ-DAF在少样本驾驶分心检测任务中显著提升性能,有效增强了模型在数据稀缺条件下的泛化能力。

原文摘要 · Abstract (English)

Driver distraction detection is essential for improving traffic safety and reducing road accidents. However, existing models often suffer from degraded generalization when deployed in real-world scenarios. This limitation primarily arises from the few-shot learning challenge caused by the high cost of data annotation in practical environments, as well as the substantial domain shift between training datasets and target deployment conditions. To address these issues, we propose a Pose-driven Quality-controlled Data Augmentation Framework (PQ-DAF) that leverages a vision-language model for sample filtering to cost-effectively expand training data and enhance cross-domain robustness. Specifically, we employ a Progressive Conditional Diffusion Model (PCDMs) to accurately capture key driver pose features and synthesize diverse training examples. A sample quality assessment module, built upon the CogVLM vision-language model, is then introduced to filter out low-quality synthetic samples based on a confidence threshold, ensuring the reliability of the augmented dataset. Extensive experiments demonstrate that PQ-DAF substantially improves performance in few-shot driver distraction detection, achieving significant gains in model generalization under data-scarce conditions.

驾驶分心数据增强少样本学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。