arXiv:2606.15072cs.CV2026-06中稿 · KDD

解决汽车近红外图像中合成数据到真实数据的分割偏差问题

Texture-Shape Bias Balancing for Robust Synthetic-to-Real Semantic Segmentation in Automotive NIR Imagery

论文配图:Texture-Shape Bias Balancing for Robust Synthetic-to-Real Semantic Segmentation in Automotive NIR Imagery
图 1 · 摘自论文原文
  • 用低秩微调扩散模型生成逼真近红外风格图像,保持结构不变
  • 引入拓扑纹理多样化策略,使分割对纹理变化更鲁棒
  • 在内外场景上分别减少63.6%和28.4%域差距,适合自动驾驶感知研究

语义分割是现代汽车视觉感知的核心,实现像素级场景理解。近红外成像(NIR)可在复杂光照下稳定检测,但因真实场景高质量标注数据稀缺,领域专用分割模型开发困难。合成数据可规模化替代,但训练于合成图像的模型在迁移到真实域时性能下降。本文首次系统研究汽车领域近红外图像中从合成到真实域的适应问题。提出一种生成增强框架,通过目标风格适配(TSA)将合成图像转化为真实近红外风格:基于少量精选真实近红外图像,用低秩适配微调潜在扩散模型,并通过结构保持的多信号条件化应用于合成数据。为降低纹理偏差、提升分割鲁棒性,进一步引入基于Voronoi的风格多样化策略(VSD),在保留场景几何的前提下修改原始纹理。在车内外场景的多模型实验表明,训练中平衡归纳偏置显著提升分割鲁棒性,在外景和内景数据上分别减少63.6%和28.4%的域差距。代码已开源。

原文摘要 · Abstract (English)

Semantic segmentation is a fundamental component of visual perception in modern automotive systems, enabling pixel-level scene understanding. Near-Infrared imaging (NIR) offers stable detection under difficult illumination conditions, but the development of domain-specific semantic segmentation models remains challenging due to the lack of high-quality annotated data from real-world scenarios. Synthetic datasets offer a scalable alternative, but models trained on synthetic images often suffer performance degradation when transferred to real domains. We present the first systematic study on synthetic to real domain adaptation for semantic segmentation in NIR images in the automotive domain. We propose a generative augmentation framework that transforms synthetic images into realistic NIR-style variants via our introduced target style adaptation (TSA). TSA fine-tunes a latent diffusion model via low-rank adaptation on a small curated set of real NIR images and applies it to synthetic training data using structure-preserving multi-signal conditioning. To reduce texture bias and improve segmentation robustness, we further apply a Voronoi-based style diversification strategy (VSD) that modifies the original textures while preserving scene geometry. Experiments with multiple model architectures on NIR data from vehicle interiors and street scenes show that balancing inductive bias during training leads to noticeably more robust semantic segmentation and effectively reduces the domain gap in our real-world scenarios by up to 63.6% on exterior and 28.4% on interior data. The code is available at GitHub.

语义分割域自适应近红外生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。