用自然语言生成复杂交通交互,让虚拟驾驶更真实。
OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space
- 通过视觉语言模型控制时空体素动态,无需轨迹等几何先验。
- 在85000条多层级语言指令上实现高保真交互生成,长时序一致性优秀。
- 适合自动驾驶仿真、行为建模研究者,推动从画面生成到行为编排的变革。
生成式世界模型越来越依赖4D占用空间进行真实的自动驾驶模拟。然而,现有生成框架依赖刚性的几何条件(如显式轨迹)或简单的属性级文本,难以编排复杂的、时序性的多智能体交互。为弥合这一语义-时空鸿沟,我们提出OccDirector,一个开创性框架,仅凭自然语言即可生成4D占用动态。作为“场景导演”,OccDirector将自然语言脚本映射为物理上合理的体素动态,无需几何先验。技术上,它采用基于视觉语言模型的时空MMDiT,并结合历史前缀锚定策略,确保长时序交互一致性。此外,我们构建了全新的数据集OccInteract-85k,首次标注了从静态布局到复杂多智能体行为的多层级语言指令,并设计了基于VLM的新评估基准。大量实验表明,OccDirector在生成质量与指令遵循能力上达到当前最佳,成功推动范式从外观合成转向语言驱动的行为编排。
原文摘要 · Abstract (English)
Generative world models increasingly rely on 4D occupancy for realistic autonomous driving simulation. However, existing generation frameworks depend on rigid geometric conditions (e.g., explicit trajectories) or simplistic attribute-level text, failing to orchestrate complex, sequential multi-agent interactions. To address this semantic-spatiotemporal gap, we propose OccDirector, a pioneering framework that generates 4D occupancy dynamics conditioned solely on natural language. Operating as a ``scenario director'', OccDirector maps natural language scripts into physically plausible voxel dynamics without requiring geometric priors. Technically, it employs a VLM-driven Spatio-Temporal MMDiT equipped with a history-prefix anchoring strategy to ensure long-horizon interaction consistency. Furthermore, we introduce OccInteract-85k, a novel dataset uniquely annotated with multi-level language instructions: ranging from static layouts to intricate multi-agent behaviors, alongside a novel VLM-based evaluation benchmark. Extensive experiments demonstrate that OccDirector achieves state-of-the-art generation quality and unprecedented instruction-following capabilities, successfully shifting the paradigm from appearance synthesis to language-driven behavior orchestration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。