用扩散模型生成风格化3D场景,稳定且可控。
DReSG: Diffusion Residuals for Stylized Gaussian Splatting

- 将扩散模型的风格提示作为残差反馈到3D高斯场景中
- 多视角融合确保跨视图一致性,避免画面漂移
- 适合需要高质量、一致风格化3D内容的创作者
基于参考图像的3D高斯点云风格化对高效可控的3D内容创作至关重要。现有基于VGG特征的方法虽能稳定优化渲染视图,但常忽略参考风格的表达细节;扩散模型具有更强的图像先验,但直接使用逐视图或基于分数的扩散引导会导致视图漂移、局部伪影和难以控制的外观更新。我们提出DReSG,一种面向风格化高斯点云的3D感知残差反馈框架。DReSG将注意力引导的扩散提议表示为相对于当前渲染的残差目标,并通过多视角高斯反馈逐步将这些残差融入共享的高斯场景中。为确保反馈的稳定性与可控性,DReSG在目标构建中调节残差强度,并在多视角拟合中结合覆盖感知的视图选择与冲突过滤的颜色更新。大量实验表明,DReSG在保持场景结构和跨视图稳定性方面表现更优,同时实现具有竞争力的参考引导风格化效果。项目页面见:https://vpx-ecnu.github.io/DReSG-website/
原文摘要 · Abstract (English)
Reference-guided stylization of scenes represented by 3D Gaussian Splatting (3DGS) is important for efficient and controllable 3D content creation. Existing VGG-feature-based 3D stylization methods provide stable rendered-view optimization, but often under-represent expressive reference style cues; diffusion models offer stronger image priors, yet direct per-view or score-based diffusion guidance can lead to view drift, local artifacts, and hard-to-control appearance updates. We present DReSG, a 3D-grounded residual-feedback framework for stylized Gaussian splatting. DReSG represents attention-guided diffusion proposals as residual targets relative to the current render, and progressively absorbs these residuals into a shared Gaussian scene through multi-view Gaussian feedback. To make this feedback stable and controllable, DReSG modulates residual strength during target construction and combines coverage-aware view selection with conflict-filtered color updates during multi-view fitting. Extensive experiments demonstrate that DReSG achieves competitive reference-guided stylization while better preserving scene structure and cross-view stability. Our project page is available at https://vpx-ecnu.github.io/DReSG-website/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。