用文本生成交通场景图像,更准更细更真实。
Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes
- 用可控制的掩码机制统一生成与编辑任务
- 结合多视角数据提升场景几何多样性
- 针对小物体设计加权损失,提升细节质量
随着智能交通系统快速发展,文本驱动的图像生成与编辑技术在交通监控、自动驾驶等应用中展现出巨大潜力。然而,现有方法仍面临生成元素语义贫乏、视角单一、视觉保真度低及图文对齐差等问题。为此,本文提出一个统一的文本驱动框架,通过可控掩码机制无缝融合图像生成与编辑;引入车端与路侧多视角数据以增强场景几何多样性;采用两阶段训练策略:先利用大规模粗粒度图文数据进行概念学习,再通过细粒度描述数据微调以提升图文对齐与细节质量;此外,提出掩码区域加权损失,在训练中动态强化小而关键区域的关注,显著提升小尺度交通元素的生成保真度。大量实验表明,该方法在交通场景的文本图像生成与编辑任务中达到领先性能。
原文摘要 · Abstract (English)
With the rapid advancement of intelligent transportation systems, text-driven image generation and editing techniques have demonstrated significant potential in providing rich, controllable visual scene data for applications such as traffic monitoring and autonomous driving. However, several challenges remain, including insufficient semantic richness of generated traffic elements, limited camera viewpoints, low visual fidelity of synthesized images, and poor alignment between textual descriptions and generated content. To address these issues, we propose a unified text-driven framework for both image generation and editing, leveraging a controllable mask mechanism to seamlessly integrate the two tasks. Furthermore, we incorporate both vehicle-side and roadside multi-view data to enhance the geometric diversity of traffic scenes. Our training strategy follows a two-stage paradigm: first, we perform conceptual learning using large-scale coarse-grained text-image data; then, we fine-tune with fine-grained descriptive data to enhance text-image alignment and detail quality. Additionally, we introduce a mask-region-weighted loss that dynamically emphasizes small yet critical regions during training, thereby substantially enhancing the generation fidelity of small-scale traffic elements. Extensive experiments demonstrate that our method achieves leading performance in text-based image generation and editing within traffic scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。