arXiv:2512.17445cs.CV2025-12中稿 · ECCV被引 3

用自然语言编辑真实驾驶视频,让交通场景随指令自由变换。

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

  • 将视频转为3D场景图,分静态背景与动态物体节点。
  • 指令对齐率超前代模型近2倍,保持真实感与交通逻辑。
  • 适合自动驾驶仿真、影视制作等需要精准场景控制的领域。

LangDriveCTRL 是一个支持自然语言控制的驾驶视频编辑框架,可合成多样化的交通场景。它将视频显式表示为3D场景图,分解为静态背景和动态物体节点。为实现精细编辑与真实感,引入反馈驱动的代理流水线:统筹者将用户指令转化为可执行图,协调多模态代理与工具;物体定位代理将自由文本与场景图中的目标节点对齐;行为编辑代理从语言指令生成多物体轨迹;行为评审代理迭代审查并优化轨迹。编辑后的场景图通过视频扩散工具渲染并融合,再由视频评审代理进一步优化,确保视觉真实与外观一致。该方法支持物体节点的增删改及多物体行为的自然语言编辑。定量评估显示,其指令对齐度接近前代最优模型的2倍,且在真实感、结构保真度与交通合理性方面表现更优。

原文摘要 · Abstract (English)

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into a static background and dynamic object nodes. To enable fine-grained editing and realism, it introduces a feedback-driven agentic pipeline. An Orchestrator converts user instructions into executable graphs that coordinate specialized multi-modal agents and tools. An Object Grounding Agent aligns free-form text with target object nodes in the scene graph; a Behavior Editing Agent generates multi-object trajectories from language instructions; and a Behavior Reviewer Agent iteratively reviews and refines the generated trajectories. The edited scene graph is rendered and harmonized using a video diffusion tool, and then further refined by a Video Reviewer Agent to ensure photorealism and appearance alignment. LangDriveCTRL supports both object node editing (removal, insertion, and replacement) and multi-object behavior editing from natural-language instructions. Quantitatively, it achieves nearly $2\times$ higher instruction alignment than the previous SoTA, with superior photorealism, structural preservation, and traffic realism. Project page is available at: https://yunhe24.github.io/langdrivectrl/.

驾驶模拟自然语言控制多模态代理视频编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。