arXiv:2606.29097cs.CV2026-06中稿 · CVPR

用真实驾驶视频训练大模型,生成更逼真的交通场景。

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

论文配图:TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
图 1 · 摘自论文原文
  • 基于真实视频合成交通场景,自动对齐大模型与现实分布。
  • 生成场景使自动驾驶模型碰撞率平均提升10.8%。
  • 适合自动驾驶测试、模型训练与跨区域交通研究。

近期研究探索使用大语言模型(LLMs)生成自动驾驶用交通场景,但预训练模型常与真实交通分布不一致。本文提出TrafficAlign,一个自动化框架,基于真实驾驶视频生成交通场景,进行数据验证,并对齐大模型与生成场景。评估显示,由TrafficAlign生成的场景在三个自动驾驶模型上平均多揭示10.8%的碰撞。此外,使用这些场景微调模型后,碰撞率相比原模型降低36.1%。六个多地理区域的定性研究表明,生成场景与各地交通分布高度一致。

原文摘要 · Abstract (English)

Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs often fail to align with real-world traffic distributions. In this work, we present TrafficAlign, an automated framework that synthesizes traffic scenarios based on real-world driving videos, performs data validation, and aligns LLMs with the synthesized scenarios. The evaluation shows that traffic scenarios generated by TrafficAlign are highly effective, revealing up to 10.8% more collisions on average across three autonomous driving models than state-of-the-art methods. Furthermore, fine-tuning these driving models with TrafficAlign-generated scenarios significantly reduced collision rates by 36.1% compared with the original models. A qualitative study using traffic datasets from six geographically diverse regions shows that TrafficAlign-generated scenarios exhibit strong alignment with corresponding traffic distributions in these regions.

交通生成大模型自动驾驶数据对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。