通过局部语义对齐提升交通视频生成的时序一致性,无需推理时控制信号。
LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
- 在动态物体附近比对真实与生成视频的语义特征,构建一致性损失
- 单轮微调即超越基线,在nuScenes和KITTI上取得更高mAP与mIoU
- 适用于自动驾驶数据生成,无需额外控制或计算开销
可控视频生成已成为自动驾驶中生成逼真交通场景的有力工具。然而,现有方法依赖推理时的控制信号引导生成模型实现动态物体的时序一致生成,限制了其作为可扩展、通用数据引擎的潜力。本文提出局部语义对齐(LSA)框架,用于微调预训练视频生成模型。LSA通过比对真实与生成视频片段在动态物体附近的语义特征,引入语义特征一致性损失,并与标准扩散损失结合进行模型微调。仅用一个周期的微调,该模型在常见视频生成评估指标上超越基线。为更深入测试时序一致性,我们借鉴目标检测任务中的mAP与mIoU两个指标。在nuScenes和KITTI数据集上的大量实验表明,本方法可在无需推理时外部控制信号及额外计算开销的前提下,有效提升视频生成的时序一致性。
原文摘要 · Abstract (English)
Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model towards temporally consistent generation of dynamic objects, limiting their utility as scalable and generalizable data engines. In this work, we propose Localized Semantic Alignment (LSA), a simple yet effective framework for fine-tuning pre-trained video generation models. LSA enhances temporal consistency by aligning semantic features between ground-truth and generated video clips. Specifically, we compare the output of an off-the-shelf feature extraction model between the ground-truth and generated video clips localized around dynamic objects inducing a semantic feature consistency loss. We fine-tune the base model by combining this loss with the standard diffusion loss. The model fine-tuned for a single epoch with our novel loss outperforms the baselines in common video generation evaluation metrics. To further test the temporal consistency in generated videos we adapt two additional metrics from object detection task, namely mAP and mIoU. Extensive experiments on nuScenes and KITTI datasets show the effectiveness of our approach in enhancing temporal consistency in video generation without the need for external control signals during inference and any computational overheads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。