用语义校准提升可见光转红外图像的准确性与一致性。
SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation

- 引入语义标签校准去噪网络中的自注意力机制,增强类别内关联。
- 生成图像在物体位置、形状和布局上更接近真实红外图像。
- 适合需要高质量合成红外数据的检测任务研究者使用。
可见光到红外图像转换为利用丰富的可见图像扩充红外训练数据提供了实用方法。扩散模型因其强大的生成能力而成为该任务的有力候选。然而,现有基于扩散的方法通常仅将语义先验作为外部条件,未显式调节去噪网络内的标记交互,导致难以保持物体位置、形状和语义布局,影响标注重用的可靠性。我们提出SC-Diff,一种语义校准的潜在扩散框架,同时利用语义先验进行条件引导和内部自注意力校准。预训练SAM3模型通过预设文本提示从可见图像中提取类别特定的语义掩码,合并为语义图并与可见图像融合作为输入条件。该图被转化为标记级语义标签,用于校准去噪网络中的自注意力。基于这些标签,我们提出语义引导的自注意力校准(SGSC),自适应地为同类别查询-键对施加正偏置。查询层面的校准强度取决于注意力在语义类别间的分散程度以及分配给查询自身类别的注意力。原始注意力得分进一步调制该偏置,使具有更强响应的同类别键获得更大校准。这种软校准在减少跨类别干扰的同时保留全局上下文交互,从而提升生成红外图像的语义一致性。大量实验表明,SC-Diff在感知质量上表现更优,并生成更有效的下游红外目标检测合成训练数据。
原文摘要 · Abstract (English)
Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because of their strong generative performance. However, existing diffusion-based methods typically use semantic priors only as external conditions, without explicitly regulating token interactions within the denoising network. Consequently, they struggle to preserve object locations, shapes, and semantic layouts required for reliable annotation reuse. We propose SC-Diff, a semantically calibrated latent diffusion framework that uses semantic priors for both conditional guidance and internal self-attention calibration. A pretrained SAM3 model with predefined text prompts first extracts category-specific semantic masks from visible images. These masks are merged into a semantic map and fused with the visible image as the input condition. The same map is converted into token-level semantic labels to calibrate self-attention in the denoising network. Based on these labels, we introduce Semantic-Guided Self-Attention Calibration (SGSC), which adaptively applies positive biases to query-key pairs of the same category. The query-wise calibration strength depends on the dispersion of attention across semantic categories and the attention assigned to the query's own category. The original attention scores further modulate the bias, giving greater calibration to same-category keys with stronger responses. This soft calibration reduces cross-category interference while retaining global contextual interactions, thereby improving semantic consistency in generated infrared images. Extensive experiments show that SC-Diff improves perceptual quality and produces more effective synthetic training data for downstream infrared object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。