将真实道路视频转为可调控的自动驾驶测试场景,提升验证真实性。
Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems

- 用视觉语言模型从视频提取语义,匹配封闭场地地图
- 生成多样多车轨迹,风险等级和交互风格可调
- 适合自动驾驶安全验证团队使用,提升测试多样性
封闭场地测试在自动驾驶系统(ADS)的验证中至关重要,尤其针对安全关键场景,可在受控条件下实现可复现评估。然而,现有方法大多依赖标准化协议或预设轨迹,导致交互过于剧本化,难以还原公共道路交通的真实复杂性。为此,我们提出 Video2Track 框架,将真实世界驾驶视频中的交互场景转化为可调控的对抗性封闭场地测试。该框架包含两个紧密耦合模块:第一是场景语义映射模块,利用视觉语言模型从驾驶视频中提取结构化语义,并通过检索增强生成技术将其定位到封闭场地拓扑库,识别兼容的地图片段与交互锚点;第二是动态交互测试模块,基于定位后的拓扑与锚点,通过条件扩散模型生成多样化的多智能体轨迹,同时通过参数化对抗目标的斯塔克尔伯格博弈调节交互强度。封闭场地实验表明,该框架能忠实复现典型真实世界交互场景,并生成具备可控风险水平与交互风格的可执行场景变体,为自动驾驶系统提供一种可扩展、可调控的现实验证方法。
原文摘要 · Abstract (English)
Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches still rely on standardized protocols or predefined trajectories, leading to overly scripted interactions and limited ability to reproduce the natural complexity of public-road traffic. To address this limitation, we propose Video2Track, a framework that transfers real-world interactive driving scenarios from videos into steerable adversarial closed-track testing. The framework consists of two tightly coupled modules. The first is a scenario semantic mapping module, which extracts structured semantics from driving videos using a vision-language model and grounds them onto a closed-track topology library via retrieval-augmented generation, thereby identifying compatible map segments and interaction anchors. The second is a dynamic interactive testing module, which conditions on the grounded topology and anchors to generate diverse multi-agent trajectories through a conditional diffusion model, while regulating interaction intensity via a Stackelberg game with a parameterized adversarial objective. Closed-track experiments demonstrate that the proposed framework can faithfully reproduce representative real-world interaction scenarios and generate executable scenario variants with controllable risk levels and interaction styles, providing a scalable approach for realistic and steerable ADS validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。