用小波相位扩散实现逼真且结构一致的仿真到现实图像转换
Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation

- 在双树复小波包变换域注入相位,实现局部自适应控制
- 低频随机化摆脱合成光照先验,提升真实感
- 无需配对数据,支持物体级独立翻译,适合自动驾驶场景
仿真到现实的图像转换需弥合合成与真实域间的外观差异,同时保持结构和语义一致性。基于条件的方法虽实现空间对齐,但引入计算开销大的控制模块;成对数据方法虽真实感强,却依赖复杂合成流程,常改变场景几何与语义;无训练编辑方法避免上述约束,但缺乏学习到的外观先验,感知质量受限。近期提出的相位保持扩散模型具前景,但傅里叶域表征受全局频谱耦合限制,导致环形伪影和边界泄漏,损害结构与语义一致性。本文提出小波相位扩散(Wavelet Phase Diffusion),通过两部分改进:其一,在双树复小波包变换域操作,局部化的小波包实现无全局频谱干扰的空间自适应相位注入;其二,低频随机化(LFR)替换低频包,解耦模型与合成光照先验,实现分布内真实世界外观。两者均在无配对开放域数据上训练,推理开销极低。空间局部性还支持实例级翻译,可独立将单个物体或区域转为逼真外观而周围场景保持不变。在vKITTI→KITTI图像翻译任务中,本方法在真实感和语义一致性上优于现有方法,同时保持良好结构对齐。在CARLA视频翻译中,接近成对数据方法的真实感,同时使视觉语言规划器的ADE和FDE分别降低5.4%和5.1%。
原文摘要 · Abstract (English)
Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism but rely on complex synthesis pipelines, often altering scene geometry and semantics. Training-free editing methods avoid both constraints but lack a learned appearance prior, limiting their perceptual quality. Recently proposed phase-preserving diffusion presents a promising alternative, but Fourier-domain formulations are constrained by global spectral coupling. This coupling induces spatial artifacts such as ringing and boundary leakage, thereby degrading structural and semantic consistency. We introduce Wavelet Phase Diffusion, which addresses this through two components. First, we operate in the Dual-Tree Complex Wavelet Packet Transform domain, whose localized wavelet packets enable spatially adaptive phase injection without global spectral interference. Second, Low-Frequency Randomization (LFR) replaces the low-frequency packet, decoupling the model from the synthetic illumination prior and enabling in-distribution real-world appearance. Both components train on unpaired open-domain data, and introduce negligible inference overhead. The spatial locality further enables instance-level translation, where individual objects or regions are translated to photorealistic appearance independently while the surrounding scene remains untranslated. On vKITTI $\to$ KITTI image translation, ours outperforms prior methods in realism and semantic consistency while maintaining competitive structural alignment. For CARLA video translation, ours approaches the realism of paired-data methods while reducing VLM planner ADE and FDE by $5.4\%$ and $5.1\%$, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。