arXiv:2602.22549cs.CVcs.AI2026-02被引 1

提升自动驾驶场景生成的细节与可控性,解决传统方法生成失败问题。

DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation

  • 分阶段学习+互信息约束,缓解几何条件间的依赖冲突。
  • 多视角语义描述增强文本引导,细化背景与结构细节。
  • 频域引导损失强化高频率结构感知,减少模糊与失真。

驾驶场景的多样化合成是验证自动驾驶系统鲁棒性与泛化能力的关键数据增强技术。现有方法在扩散模型中以高清地图和3D边界框作为几何条件进行条件生成,但隐式条件间依赖导致控制条件独立变化时生成失败。同时,语义与结构细节不足:简短且视角不变的描述限制了语义上下文,造成背景建模薄弱;标准去噪损失采用均匀空间加权,忽略前景结构细节,引发视觉失真与模糊。为此,我们提出DrivePTS,包含三项创新:首先,采用分阶段学习策略并引入显式互信息约束,缓解几何条件间的依赖;其次,利用视觉-语言模型生成跨六种语义维度的多视角层级描述,提供细粒度文本指导;第三,引入频域引导结构损失,增强模型对高频元素的敏感性,提升前景结构保真度。大量实验表明,DrivePTS在生成多样驾驶场景方面达到当前最优的保真度与可控性,尤其能成功生成先前方法失效的罕见场景,凸显其强大泛化能力。

原文摘要 · Abstract (English)

Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes as geometric conditions in diffusion models for conditional scene generation. However, implicit inter-condition dependency causes generation failures when control conditions change independently. Additionally, these methods suffer from insufficient details in both semantic and structural aspects. Specifically, brief and view-invariant captions restrict semantic contexts, resulting in weak background modeling. Meanwhile, the standard denoising loss with uniform spatial weighting neglects foreground structural details, causing visual distortions and blurriness. To address these challenges, we propose DrivePTS, which incorporates three key innovations. Firstly, our framework adopts a progressive learning strategy to mitigate inter-dependency between geometric conditions, reinforced by an explicit mutual information constraint. Secondly, a Vision-Language Model is utilized to generate multi-view hierarchical descriptions across six semantic aspects, providing fine-grained textual guidance. Thirdly, a frequency-guided structure loss is introduced to strengthen the model's sensitivity to high-frequency elements, improving foreground structural fidelity. Extensive experiments demonstrate that our DrivePTS achieves state-of-the-art fidelity and controllability in generating diverse driving scenes. Notably, DrivePTS successfully generates rare scenes where prior methods fail, highlighting its strong generalization ability.

场景生成扩散模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。