统一修复扩散模型的视图合成失真问题
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis

- 通过参考图像粗细结合修正退化,提升生成质量
- 在多个数据集上实现先进性能,支持零样本适配
- 适合需要高质量3D视觉生成的研究者与开发者
随着生成模型的兴起,基于扩散的视图合成方法已成为主流,无论是显式深度扭曲重绘还是隐式端到端方式。然而,两类方法常因像素到潜在空间压缩及扩散幻觉导致细节模糊、结构扭曲等质量下降问题。本文从空间、时间与骨干网络三个维度分析扩散退化,并提出UniFixer——一种通用的参考引导修复框架,采用粗到精策略:首先设计参考预对齐模块实现参考视图与退化新视图的粗略对齐;随后全局结构锚定机制修正几何畸变以保证结构保真;最后局部细节注入模块恢复精细纹理,实现高质量视图合成。UniFixer作为即插即用的后处理模块,可在不同类型的扩散退化中实现零样本修复,大量实验验证其在新视图合成与立体转换任务上的领先表现。
原文摘要 · Abstract (English)
With the recent surge of generative models, diffusion-based approaches have become mainstream for view synthesis tasks, either in an explicit depth-warp-inpaint or in an implicit end-to-end manner. Despite their success, both paradigms often suffer from noticeable quality degradation, e.g., blurred details and distorted structures, caused by pixel-to-latent compression and diffusion hallucination. In this paper, we investigate diffusion degradation from three key dimensions (i.e., spatial, temporal, and backbone-related) and propose UniFixer, a universal reference-guided framework that fixes diverse degradation artifacts via a coarse-to-fine strategy. Specifically, a reference pre-alignment module is first designed to perform coarse alignment between the reference view and the degraded novel view. A global structure anchoring mechanism then rectifies geometric distortions to ensure structural fidelity, followed by a local detail injection module that recovers fine-grained texture details for high-quality view synthesis. Our UniFixer serves as a plug-and-play refiner that achieves zero-shot fixing across different types of diffusion degradation, and extensive experiments verify our state-of-the-art performance on novel view synthesis and stereo conversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。