解决昼夜图像转换中目标物体幻觉问题,提升下游任务准确性。
Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image Translation

- 用双头判别器结合语义分割检测背景中的幻觉内容。
- 通过类别原型抑制幻觉,使目标类特征在特征空间中保持稳定。
- 在BDD100K上对易幻觉类别的检测精度提升31.7%。
昼夜无配对图像转换对下游任务至关重要,但因外观差异大且缺乏像素级监督而困难重重。现有方法常产生语义幻觉,如错误合成交通标志、车辆及人工光源等目标类内容,显著降低下游性能。本文提出新框架,在无配对翻译过程中检测并抑制目标类特征的幻觉。为检测幻觉,设计双头判别器,额外执行语义分割以识别背景中的幻觉区域;为抑制幻觉,引入类别特定原型,通过聚合标注目标域物体特征构建,作为各类别的语义锚点。基于薛定谔桥的翻译模型实现迭代优化,将检测到的幻觉特征在特征空间中主动远离类别原型,从而保持跨翻译轨迹的物体语义一致性。实验表明,本方法在定性和定量上均优于现有方法。在BDD100K数据集上,日到夜域自适应的mAP提升15.5%,对易幻觉类别(如交通灯)的提升达31.7%。
原文摘要 · Abstract (English)
Day-to-night unpaired image translation is important to downstream tasks but remains challenging due to large appearance shifts and the lack of direct pixel-level supervision. Existing methods often introduce semantic hallucinations, where objects from target classes such as traffic signs and vehicles, as well as man-made light effects, are incorrectly synthesized. These hallucinations significantly degrade downstream performance. We propose a novel framework that detects and suppresses hallucinations of target-class features during unpaired translation. To detect hallucination, we design a dual-head discriminator that additionally performs semantic segmentation to identify hallucinated content in background regions. To suppress these hallucinations, we introduce class-specific prototypes, constructed by aggregating features of annotated target-domain objects, which act as semantic anchors for each class. Built upon a Schrodinger Bridge-based translation model, our framework performs iterative refinement, where detected hallucination features are explicitly pushed away from class prototypes in feature space, thus preserving object semantics across the translation trajectory.Experiments show that our method outperforms existing approaches both qualitatively and quantitatively. On the BDD100K dataset, it improves mAP by 15.5% for day-to-night domain adaptation, with a notable 31.7% gain for classes such as traffic lights that are prone to hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。