零样本生成驾驶场景,无需微调即可实现多视角一致与恶劣天气泛化。
FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model

- 采用无参数冻结扩散模型,通过时空注意力实现跨视角对齐和时间连贯性。
- 在nuScenes数据集上,夜间与雨天场景性能提升显著,罕见类别生成更精准。
- 适合自动驾驶数据增强、罕见场景合成,无需针对特定天气微调。
自动驾驶的合成数据正迅速增长,扩散模型推动了可扩展的场景生成。然而,多视角与时间一致性通常依赖主干网络微调或附加层,破坏预训练知识并削弱文本对齐能力。现有模型仍局限于训练分布,在极端天气和未知配置下表现不佳,且生成质量偏向高频类别。我们提出FrozenDrive,一个可控生成框架,在保持预训练扩散模型知识的同时实现强一致性。该方法基于丰富的驾驶栈信号与文本提示,引入知识保持型时空注意力,在单次前向传播中实现跨视角对齐与时间连贯性;额外的物体聚焦约束提升了稀有类别的局部保真度。无需任何天气或场景特化微调,模型可从文本生成全局一致的多视角驾驶场景,尤其在恶劣与罕见条件下表现优异。在nuScenes数据集上,使用FrozenDrive生成的数据显著提升自动驾驶模型性能,尤其在夜间与雨天场景,验证了其在场景定向数据训练下的更强鲁棒性。
原文摘要 · Abstract (English)
Synthetic data for autonomous driving is surging, powered by diffusion models that promise scalable scene generation. Yet key obstacles remain, as enforcing multi-view and temporal consistency often relies on backbone fine-tuning or added layers, which erodes pre-trained knowledge and weakens text alignment. Models also stay close to the training distribution, struggling under adverse weather and unseen configurations, and fidelity favors frequent over rare classes. We address these gaps with FrozenDrive, a controllable generative framework that preserves a pretrained diffusion models knowledge while achieving strong consistency. FrozenDrive conditions on rich driving-stack signals and text prompts, and introduces knowledge-preserving spatio-temporal attention to impose cross-view alignment and temporal coherence in a single pass within a parameter-free frozen diffusion backbone. An additional object-focused constraint improves per-object fidelity for rare categories. Without any weather- or scene-specific fine-tuning, our model synthesizes globally coherent multi-view driving scenes from text, particularly under adverse and rare conditions, and surpasses prior baselines. On nuScenes, FrozenDrive augmented data significantly improves AD models performance, especially at night and in rain, demonstrating stronger robustness when trained with our scenario-targeted data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。