AI生成的驾驶视频能用于自动驾驶训练吗?这份诊断框架给出了答案。
Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- 构建故障分类体系,识别生成视频中的视觉瑕疵与物理错误
- 实测表明未经筛选的AI视频会降低感知模型性能
- 提出多维度评分器ADGVE,可有效筛选高质量生成视频
近期文本到视频模型已能从自然语言提示生成高分辨率驾驶场景。这些AI生成驾驶视频(AIGVs)为自动驾驶(AD)提供了低成本、可扩展的数据替代方案。但关键问题仍存:它们能否可靠地支持AD模型的训练与评估?本文提出一个诊断框架,系统研究该问题。首先,我们建立常见AIGV失效模式的分类体系,包括视觉伪影、物理上不合理的运动及交通语义违规,并验证其对目标检测、跟踪和实例分割的负面影响。为支持分析,我们构建了专注于驾驶任务的基准ADGV-Bench,包含人工质量标注和多个感知任务的密集标签。随后提出ADGVE——一种驾驶感知评估器,融合静态语义、时序线索、车道合规信号与视觉-语言模型(VLM)引导推理,生成每段视频的综合质量分数。实验显示,直接使用原始AIGVs会损害感知性能,而通过ADGVE筛选后,不仅提升视频质量评估指标,还显著改善下游AD模型表现,使AIGVs成为真实数据的有益补充。本研究揭示了AIGVs的潜在风险与价值,并为未来自动驾驶流水线中安全利用大规模视频生成技术提供实用工具。
原文摘要 · Abstract (English)
Recent text-to-video models have enabled the generation of high-resolution driving scenes from natural language prompts. These AI-generated driving videos (AIGVs) offer a low-cost, scalable alternative to real or simulator data for autonomous driving (AD). But a key question remains: can such videos reliably support training and evaluation of AD models? We present a diagnostic framework that systematically studies this question. First, we introduce a taxonomy of frequent AIGV failure modes, including visual artifacts, physically implausible motion, and violations of traffic semantics, and demonstrate their negative impact on object detection, tracking, and instance segmentation. To support this analysis, we build ADGV-Bench, a driving-focused benchmark with human quality annotations and dense labels for multiple perception tasks. We then propose ADGVE, a driving-aware evaluator that combines static semantics, temporal cues, lane obedience signals, and Vision-Language Model(VLM)-guided reasoning into a single quality score for each clip. Experiments show that blindly adding raw AIGVs can degrade perception performance, while filtering them with ADGVE consistently improves both general video quality assessment metrics and downstream AD models, and turns AIGVs into a beneficial complement to real-world data. Our study highlights both the risks and the promise of AIGVs, and provides practical tools for safely leveraging large-scale video generation in future AD pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。