arXiv:2602.09476cs.CV2026-02

解决合成到真实图像翻译中逼真度与结构稳定性的矛盾

FD-DB: Frequency-Decoupled Dual-Branch Network for Unpaired Synthetic-to-Real Domain Translation

  • 分频双分支架构,低频做可解释编辑,高频补细节
  • 在YCB-V数据集上提升语义分割性能,保持几何结构一致
  • 适合需要高质量合成数据的视觉任务研究者

合成数据为几何敏感的视觉任务提供低成本、精确标注样本,但合成与真实域之间的外观和成像差异导致严重域偏移,降低下游性能。无配对的合成到真实图像翻译可缓解这一问题,但现有方法常面临逼真度与结构稳定性之间的权衡:自由生成可能引入形变或虚假纹理,而过度约束则限制对真实域统计特性的适应。我们提出FD-DB,一种频率解耦的双分支网络,将外观迁移分解为低频可解释编辑与高频残差补偿。可解释分支预测物理意义明确的编辑参数(白平衡、曝光、对比度、饱和度、模糊、噪点),构建稳定的低频外观基础,保证内容保真;自由分支通过残差生成补充细粒度细节,门控融合机制在显式频率约束下结合两分支,限制低频漂移。进一步采用两阶段训练策略,先稳定编辑分支,再释放残差分支以提升优化稳定性。在YCB-V数据集上的实验表明,FD-DB显著提升真实域外观一致性,并大幅改善下游语义分割性能,同时保持几何与语义结构。

原文摘要 · Abstract (English)

Synthetic data provide low-cost, accurately annotated samples for geometry-sensitive vision tasks, but appearance and imaging differences between synthetic and real domains cause severe domain shift and degrade downstream performance. Unpaired synthetic-to-real translation can reduce this gap without paired supervision, yet existing methods often face a trade-off between photorealism and structural stability: unconstrained generation may introduce deformation or spurious textures, while overly rigid constraints limit adaptation to real-domain statistics. We propose FD-DB, a frequency-decoupled dual-branch model that separates appearance transfer into low-frequency interpretable editing and high-frequency residual compensation. The interpretable branch predicts physically meaningful editing parameters (white balance, exposure, contrast, saturation, blur, and grain) to build a stable low-frequency appearance base with strong content preservation. The free branch complements fine details through residual generation, and a gated fusion mechanism combines the two branches under explicit frequency constraints to limit low-frequency drift. We further adopt a two-stage training schedule that first stabilizes the editing branch and then releases the residual branch to improve optimization stability. Experiments on the YCB-V dataset show that FD-DB improves real-domain appearance consistency and significantly boosts downstream semantic segmentation performance while preserving geometric and semantic structures.

域迁移图像生成合成数据双分支

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。