构建大规模驾驶风格数据集,解决真实路况下司机特征识别难题。
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

- 基于465名司机的975小时多模态驾驶数据,定义稳定驾驶行为模式。
- 新模型在少样本识别中达到0.935的AUROC,显著优于传统方法。
- 揭示视频模型存在路线泄露问题,强调行为真实性评估必要性。
驾驶风格反映驾驶员在相似条件下稳定的、个体化的驾驶行为模式。然而,在自然驾驶数据中,这种信号难以分离,因驾驶员在不同车辆、道路和环境下行驶,导致模型可能将车辆或情境特有规律误认为司机专属风格。本文提出DriveDNA,一个大规模自然驾驶数据集与基准测试,涵盖465名驾驶员的4,121次驾驶记录,涉及115种车型,总计975小时人类控制驾驶数据(10 Hz采样率),包含前向视频。该数据集将驾驶风格定义为在相似条件下具有一致性的驾驶员行为模式。基准测试通过三项核心任务评估:少样本驾驶员再识别、个性化行为预测及条件匹配比较,并提供行为标注与276,248个规则生成的驾驶事件(六类),经大规模人工审核。我们评估了从经典特征描述符到监督/自监督时序编码器、多模态融合、概率预测及零样本基础模型等基线方法,采用固定多种子协议。学习表征在未见驾驶员上显著优于传统描述符(AUROC 0.935 vs. 0.707),且在匹配条件下保留司机特异性信息;而描述符性能接近随机水平。仅使用视频的模型虽达到可比识别准确率,但存在严重路线泄露现象,表明强识别能力可能源于上下文捷径而非真实驾驶行为。结果表明,可靠的驾驶风格评估需同时考察学习表征的行为价值及其对车辆、行程和条件混淆因素的鲁棒性。
原文摘要 · Abstract (English)
Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。