arXiv:2603.28353cs.CV2026-03被引 1

让自动驾驶视频生成更精准可控,支持物体级精细控制与长期一致性。

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

  • 通过多视角视觉语言推理实现物体级细粒度控制
  • 生成长视频时对尾部物体的可控性提升显著
  • 自研评估闭环机制保障时空一致性,适合复杂场景生成

驾驶视频生成在可控性、分辨率和长度方面已取得显著进展,但在长视频生成中仍难以支持对特定实体的细粒度对象级控制,且难以保持时空一致性。本文提出VistaGEN,一种新型驾驶视频生成技术,可在长视频序列中实现对3D物体、图像及文本描述等特定实体的细粒度控制,同时维持时空一致性。其核心创新在于将多视角视觉语言推理引入长视频生成流程:通过向多视角视频生成器注入视觉语言特征以实现细粒度控制;并设计多视角视觉语言评估器(MV-VLM),智能自动评估生成内容的时空一致性,构建“生成-评估-重生成”的闭环机制,确保高质量、连贯输出,支持复杂可靠驾驶场景的生成。此外,在闭环中引入对象级精修模块,对MV-VLM评估不满意的生成结果进行优化后反馈至生成器重生成。大量实验表明,VistaGEN在长尾物体控制能力上显著优于以往方法,时空一致性大幅提升。

原文摘要 · Abstract (English)

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatiotemporal consistency, especially in long video generation. In this paper, we present a new driving video generation technique, called VistaGEN, which enables fine-grained control of specific entities, including 3D objects, images, and text descriptions, while maintaining spatiotemporal consistency in long video sequences. Our key innovation is the incorporation of multiview visual-language reasoning into the long driving video generation. To this end, we inject visual-language features into a multiview video generator to enable fine-grained controllability. More importantly, we propose a multiview vision-language evaluator (MV-VLM) to intelligently and automatically evaluate spatiotemporal consistency of the generated content, thus formulating a novel generation-evaluation-regeneration closed-loop generation mechanism. This mechanism ensures high-quality, coherent outputs, facilitating the creation of complex and reliable driving scenarios. Besides, within the closed-loop generation, we introduce an object-level refinement module to refine the unsatisfied results evaluated from the MV-VLM and then feed them back to the video generator for regeneration. Extensive evaluation shows that our VistaGEN achieves diverse driving video generation results with fine-grained controllability, especially for long-tail objects, and much better spatiotemporal consistency than previous approaches.

视频生成自动驾驶多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。