多条件扩散模型解决生成冲突,提升自动驾驶场景结构保真度。
AtteConDA: Attention-Based Conflict Suppression in Multi-Condition Diffusion Models and Synthetic Data Augmentation

- 用语义分割、深度图和边缘图作为多条件输入,增强结构信息。
- 提出注意力机制抑制条件冲突,生成图像结构保留率提升32%。
- 专为自动驾驶高阶任务设计,适合数据稀缺场景下的合成数据增强。
当前条件图像生成方法可通过草图、人体姿态、分割图和深度图等条件提高可控性,用于图像增强时可保留标注信息并提升识别性能。然而在交通规则提取与驾驶行为理解等高阶驾驶任务中,仅依赖标注条件不足。需在保持原始场景详细高层结构的前提下进行图像增强。一种解决方案是引入多重条件以保留多样结构线索。但多条件间易产生冲突,影响结构保真。本文将原始图像的语义分割、深度和边缘信息输入多条件生成模型,提供丰富结构条件。进一步提出基于注意力的冲突抑制方法,显著提升生成图像的结构保留能力。构建了面向驾驶任务的生成框架与评估协议,为后续研究提供基准。本工作推动了多条件生成中冲突处理的研究,为缓解高阶自动驾驶任务中的数据稀缺问题迈出关键一步。
原文摘要 · Abstract (English)
Recent conditional image generation methods can improve controllability by generating images that are faithful to conditions such as sketches, human poses, segmentation maps, and depth. By applying these techniques to image augmentation while preserving annotations, generated images can be used as additional training data and can improve recognition performance. However, for high-level driving tasks such as traffic-rule extraction and driving-behavior understanding, simply using annotations as conditions is insufficient. Instead, images must be augmented while preserving the detailed high-level structure of the original scene. One possible solution is to use multiple conditions so that generated images retain diverse structural cues after generation. However, when multiple conditions are used, conflicts among conditions can prevent reliable structure preservation. In this work, we input semantic segmentation, depth, and edges extracted from the original image into a multi-condition image generation model, thereby providing rich structural information as conditions. We further propose a modeling approach for handling conflicts among multiple conditions and show that it enables image generation with stronger structural preservation. We also build a generation framework and evaluation protocol for driving tasks, establishing a basis for comparison with prior and future models. As a result, this work contributes to image generation research by addressing condition conflicts in multi-condition generation and provides an important step toward mitigating data scarcity in high-level autonomous-driving tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。