arXiv:2502.10377cs.CVcs.GR2025-02被引 1

单图风格迁移至多视角真实场景,保持语义一致性和3D一致性

ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences

  • 基于开放词汇分割建立风格图与实景图的细粒度语义对应
  • 在扩散模型中用无训练注意力机制实现单视图风格迁移
  • 通过深度引导的形变-精修网络提升多视角一致性

我们提出 ReStyle3D,一种从单张风格图像向由多视角表示的真实世界场景进行场景级外观迁移的新框架。该方法结合显式语义对应与多视角一致性,实现精确且连贯的风格化。不同于传统全局应用参考风格的方法,ReStyle3D 使用开放词汇分割建立风格图与实景图之间的密集、实例级对应关系,确保每个物体均以语义匹配的纹理进行风格化。首先,通过扩散模型中的无训练语义注意力机制将风格迁移到单个视图;随后,借助由单目深度和像素级对应关系指导的可学习形变-精修网络,将风格扩展至其他视图。实验表明,ReStyle3D 在结构保留、感知风格相似性和多视角一致性方面均优于现有方法。用户研究进一步验证其生成照片级真实感、语义忠实的结果。代码、预训练模型和数据集将公开发布,以支持室内设计、虚拟布景和3D一致风格化等新应用。

原文摘要 · Abstract (English)

We introduce ReStyle3D, a novel framework for scene-level appearance transfer from a single style image to a real-world scene represented by multiple views. The method combines explicit semantic correspondences with multi-view consistency to achieve precise and coherent stylization. Unlike conventional stylization methods that apply a reference style globally, ReStyle3D uses open-vocabulary segmentation to establish dense, instance-level correspondences between the style and real-world images. This ensures that each object is stylized with semantically matched textures. It first transfers the style to a single view using a training-free semantic-attention mechanism in a diffusion model. It then lifts the stylization to additional views via a learned warp-and-refine network guided by monocular depth and pixel-wise correspondences. Experiments show that ReStyle3D consistently outperforms prior methods in structure preservation, perceptual style similarity, and multi-view coherence. User studies further validate its ability to produce photo-realistic, semantically faithful results. Our code, pretrained models, and dataset will be publicly released, to support new applications in interior design, virtual staging, and 3D-consistent stylization.

风格迁移3D生成多视图扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。