arXiv:2604.17147cs.CVcs.RO2026-04

用文字或图片生成逼真驾驶场景,支持多视角长期延续

ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation

论文配图:ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation
图 1 · 摘自论文原文
  • 在向量隐空间中联合建模道路与动态车辆,实现统一控制
  • 跨模态控制机制让文本/图像指令精准影响路况布局和交通状态
  • 生成结果真实且时序一致,适合自动驾驶仿真测试

我们提出ScenarioControl,首个基于视觉-语言的可控制驾驶场景生成方法。给定文本提示或输入图像,该方法可合成多样且真实的3D场景回放,包含地图、随时间变化的动态物体(如车辆、行人)、基础设施及自车摄像头观测。通过在向量隐空间中联合表示道路结构与动态主体,实现了对场景的整体控制。为连接多模态输入与稀疏的向量化场景元素,我们设计了跨全局控制机制,结合交叉注意力与轻量级全局上下文分支,实现对道路布局和交通条件的细粒度调控,同时保持高真实性。该方法能从场景中不同主体视角生成时序一致的长时程场景回放。为支持训练与评估,我们发布了包含文本标注与向量地图结构对齐的数据集。大量实验表明,ScenarioControl在控制遵循度与保真度方面均优于所有对比方法。

原文摘要 · Abstract (English)

We introduce ScenarioControl, the first vision-language control mechanism for learned driving scenario generation. Given a text prompt or an input image, Scenario-Control synthesizes diverse, realistic 3D scenario rollouts - including map, 3D boxes of reactive actors over time, pedestrians, driving infrastructure, and ego camera observations. The method generates scenes in a vectorized latent space that represents road structure and dynamic agents jointly. To connect multimodal control with sparse vectorized scene elements, we propose a cross-global control mechanism that integrates crossattention with a lightweight global-context branch, enabling fine-grained control over road layout and traffic conditions while preserving realism. The method produces temporally consistent scenario rollouts from the perspectives different actors in the scene, supporting long-horizon continuation of driving scenarios. To facilitate training and evaluation, we release a dataset with text annotations aligned to vectorized map structures. Extensive experiments validate that the control adherence and fidelity of ScenarioControl compare favorable to all tested methods across all experiments. Project webpage: https://light.princeton.edu/ScenarioControl

场景生成视觉语言自动驾驶向量空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。