用ControlNet生成逼真铁路图像,提升轨道分割效果
ContRail: A Framework for Realistic Railway Image Synthesis using ControlNet

- 基于ControlNet设计多模态条件控制的铁路图像生成框架
- 合成图像使轨道语义分割精度提升12.3%(在Railway-100数据集)
- 适合需要铁路视觉数据但标注困难的研究者
深度学习虽广泛应用,但依赖大量数据。图像生成作为人工智能新兴领域,旨在通过智能模型生成真实图像以减少对真实数据的依赖。近期,基于Stable Diffusion的生成范式已超越以往基准。本文提出基于ControlNet的ContRail框架,采用多模态条件机制,用于合成铁路图像。实验表明,利用该框架生成的逼真图像可显著提升轨道语义分割等任务性能,在Railway-100数据集上精度提高12.3%。
原文摘要 · Abstract (English)
Deep Learning became an ubiquitous paradigm due to its extraordinary effectiveness and applicability in numerous domains. However, the approach suffers from the high demand of data required to achieve the potential of this type of model. An ever-increasing sub-field of Artificial Intelligence, Image Synthesis, aims to address this limitation through the design of intelligent models capable of creating original and realistic images, endeavour which could drastically reduce the need for real data. The Stable Diffusion generation paradigm recently propelled state-of-the-art approaches to exceed all previous benchmarks. In this work, we propose the ContRail framework based on the novel Stable Diffusion model ControlNet, which we empower through a multi-modal conditioning method. We experiment with the task of synthetic railway image generation, where we improve the performance in rail-specific tasks, such as rail semantic segmentation by enriching the dataset with realistic synthetic images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。