arXiv:2512.05106cs.CVcs.GR2025-12被引 5

保留图像相位信息的扩散模型,实现结构对齐的图像重渲染。

NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation

  • 通过保留输入图像相位,仅随机化幅度来构建新扩散过程。
  • 在CARLA模拟器中显著提升仿真到现实的规划器迁移性能。
  • 无需额外参数或架构改动,适用于图像与视频生成任务。

标准扩散模型使用高斯噪声破坏数据,其傅里叶系数的幅值和相位均为随机。这虽适合无条件或文本到图像生成,但会破坏空间结构,不适用于需几何一致性的任务,如重渲染、仿真增强和图像到图像转换。本文提出相位保持扩散(ϕ-PD),一种模型无关的扩散过程重构方法,在随机化幅度的同时保留输入相位,实现无需架构修改或新增参数的结构对齐生成。进一步提出频率选择性结构噪声(FSS),通过单一频率截断参数实现结构刚度的连续控制。ϕ-PD无推理开销,兼容任意图像或视频扩散模型。在真实感与风格化重渲染、驾驶规划器的仿真到现实增强任务中,ϕ-PD均产生可控且空间对齐的结果。应用于CARLA模拟器时,显著提升仿真到现实的规划器迁移性能。该方法可与现有条件控制方式互补,广泛适用于图像到图像及视频到视频生成。视频、更多示例与代码见项目页。

原文摘要 · Abstract (English)

Standard diffusion corrupts data using Gaussian noise whose Fourier coefficients have random magnitudes and random phases. While effective for unconditional or text-to-image generation, corrupting phase components destroys spatial structure, making it ill-suited for tasks requiring geometric consistency, such as re-rendering, simulation enhancement, and image-to-image translation. We introduce Phase-Preserving Diffusion (ϕ-PD), a model-agnostic reformulation of the diffusion process that preserves input phase while randomizing magnitude, enabling structure-aligned generation without architectural changes or additional parameters. We further propose Frequency-Selective Structured (FSS) noise, which provides continuous control over structural rigidity via a single frequency-cutoff parameter. ϕ-PD adds no inference-time cost and is compatible with any diffusion model for images or videos. Across photorealistic and stylized re-rendering, as well as sim-to-real enhancement for driving planners, ϕ-PD produces controllable, spatially aligned results. When applied to the CARLA simulator, ϕ-PD significantly improves sim-to-real planner transfer performance. The method is complementary to existing conditioning approaches and broadly applicable to image-to-image and video-to-video generation. Videos, additional examples, and code are available on our \href{https://yuzeng-at-tri.github.io/ppd-page/}{project page}.

扩散模型图像生成结构对齐视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。