arXiv:2510.02469cs.ROcs.AI2025-10被引 1

用语言控制生成可交互的4D驾驶场景,提升真实感与编辑精度。

SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

  • 基于语言对齐的4D高斯点云构建场景图,实现自然语言查询与编辑。
  • 多智能体路径优化使修改后行为物理合理,失败率低于基线一半。
  • 支持行人精细操作,适合自动驾驶测试与仿真场景自动构建。

利用真实传感器数据进行驾驶场景操控已成为传统模拟器的有力替代。尽管语言控制与神经场景表示取得进展,现有方法仍将定位、编辑与仿真视为松散关联的阶段,依赖启发式目标定位、人工引导和单智能体验证,限制了语义表达能力并阻碍可扩展、响应式的场景生成。我们提出SIMSplat,一种基于场景图的4D高斯点云驱动的场景编辑框架,通过在高斯场景节点中嵌入外观、运动和位置语义,使重建场景可通过自由形式自然语言查询。该框架将语言理解与对象级编辑及多智能体仿真统一于一个系统中。基于此语言接地场景图,SIMSplat支持细粒度行人操作,其多智能体路径精修模块可传播变化至所有智能体,确保反应式且物理合理的模拟。该流程还集成视觉-语言模型以实现自动化场景挖掘。实验表明,SIMSplat的接地准确率超过基线两倍,任务完成率最高,各类驾驶场景下的失败率最低。

原文摘要 · Abstract (English)

Driving scene manipulation using real-world sensor data has emerged as a promising alternative to traditional driving simulators. Despite advances in language control and neural scene representations, existing methods treat grounding, editing, and simulation as loosely connected stages, relying on heuristic object localization, manual guidance, and single-agent validation, thereby constraining semantic expressiveness and hindering scalable, reactive scenario generation. We introduce SIMSplat, a driving scene editor built on scene-graph-based 4D Gaussian Splatting augmented with language-aligned features. By embedding appearance, motion, and location semantics directly into Gaussian scene-graph nodes, SIMSplat makes reconstructed scenes queryable through free-form natural language, bridging language understanding to object-level editing and multi-agent simulation within a single framework. Building on this language-grounded scene graph, SIMSplat supports diverse edits including fine-grained pedestrian manipulation, while a multi-agent path refinement module propagates changes across all agents to ensure reactive, physically plausible simulations. The pipeline further integrates with Vision-Language Models for automated scenario mining. Experiments show that SIMSplat more than doubles baseline grounding accuracy, achieves the highest task completion rate, and produces the lowest failure rates across diverse driving scenarios.

4D高斯驾驶仿真语言对齐多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。