arXiv:2604.04887cs.CV2026-04被引 1

让AI能精准编辑复杂驾驶场景,提升自动驾驶安全测试能力。

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

  • 用语言指令控制物体与场景级编辑,支持细粒度调整。
  • 生成25.5万张图像,用户偏好提升46.4%,分割精度提高33%。
  • 适合自动驾驶仿真、场景生成与可控内容创作研究者。

自动驾驶安全性依赖于可扩展的、逼真的、可控的驾驶场景生成,远超真实世界测试范围。现有基于指令的图像编辑器因训练数据以物体为中心或艺术类为主,在密集、高安全性的驾驶布局上表现不佳。本文提出HorizonWeaver,解决三个核心挑战:(1) 多层级粒度,需在密集环境中实现物体与场景级的一致编辑;(2) 丰富高层语义,保持多样物体并遵循详细指令;(3) 普遍存在的域偏移,应对未见环境中的气候、布局与交通变化。其核心贡献包括:(1) 数据:构建基于Boreas、nuScenes和Argoverse2的配对真实/合成数据集,提升泛化能力;(2) 模型:引入语言引导掩码,通过语义增强掩码与提示实现精确语言驱动编辑;(3) 训练:采用内容保真与指令对齐联合损失,确保场景一致性与指令忠实性。整体框架实现了可扩展的逼真、指令驱动驾驶场景编辑,涵盖13类编辑任务,生成255,000张图像,在L1、CLIP、DINO指标上优于基线,用户偏好提升46.4%,BEV分割IoU提升33%。

原文摘要 · Abstract (English)

Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centric or artistic data, struggle with dense, safety-critical driving layouts. We propose HorizonWeaver, which tackles three fundamental challenges in driving scene editing: (1) multi-level granularity, requiring coherent object- and scene-level edits in dense environments; (2) rich high-level semantics, preserving diverse objects while following detailed instructions; and (3) ubiquitous domain shifts, handling changes in climate, layout, and traffic across unseen environments. The core of HorizonWeaver is a set of complementary contributions across data, model, and training: (1) Data: Large-scale dataset generation, where we build a paired real/synthetic dataset from Boreas, nuScenes, and Argoverse2 to improve generalization; (2) Model: Language-Guided Masks for fine-grained editing, where semantics-enriched masks and prompts enable precise, language-guided edits; and (3) Training: Content preservation and instruction alignment, where joint losses enforce scene consistency and instruction fidelity. Together, HorizonWeaver provides a scalable framework for photorealistic, instruction-driven editing of complex driving scenes, collecting 255K images across 13 editing categories and outperforming prior methods in L1, CLIP, and DINO metrics, achieving +46.4% user preference and improving BEV segmentation IoU by +33%. Project page: https://msoroco.github.io/horizonweaver/

自动驾驶场景编辑可控生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。