arXiv:2601.19717cs.CV2026-01被引 2

用注意力优化实现3D高斯风格迁移的多视角一致性

DiffStyle3D: Consistent 3D Gaussian Stylization via Attention Optimization

  • 在隐空间直接优化,通过注意力对齐实现风格迁移
  • 多视角一致性提升,生成内容更真实且视觉连贯
  • 适合需要高质量3D风格化的内容创作者

3D风格迁移可生成视觉表现力强的3D内容,丰富场景与物体的外观。现有基于VGG或CLIP的方法难以在模型内部建模多视角一致性,而扩散方法虽能捕捉一致性,却依赖去噪方向,导致训练不稳定。为此,我们提出DiffStyle3D,一种新型基于扩散的3DGS风格迁移范式,直接在隐空间进行优化。具体地,引入注意力感知损失,在自注意力空间对齐风格特征,同时通过内容特征对齐保留原始信息。受3D风格化几何不变性启发,提出几何引导的多视角一致性方法,将几何信息融入自注意力以建模跨视角对应关系。此外,基于几何信息构建几何感知掩码,避免视图重叠区域的冗余优化,进一步提升多视角一致性。大量实验表明,DiffStyle3D优于当前最优方法,实现更高风格化质量与视觉真实性。

原文摘要 · Abstract (English)

3D style transfer enables the creation of visually expressive 3D content, enriching the visual appearance of 3D scenes and objects. However, existing VGG- and CLIP-based methods struggle to model multi-view consistency within the model itself, while diffusion-based approaches can capture such consistency but rely on denoising directions, leading to unstable training. To address these limitations, we propose DiffStyle3D, a novel diffusion-based paradigm for 3DGS style transfer that directly optimizes in the latent space. Specifically, we introduce an Attention-Aware Loss that performs style transfer by aligning style features in the self-attention space, while preserving original content through content feature alignment. Inspired by the geometric invariance of 3D stylization, we propose a Geometry-Guided Multi-View Consistency method that integrates geometric information into self-attention to enable cross-view correspondence modeling. Based on geometric information, we additionally construct a geometry-aware mask to prevent redundant optimization in overlapping regions across views, which further improves multi-view consistency. Extensive experiments show that DiffStyle3D outperforms state-of-the-art methods, achieving higher stylization quality and visual realism.

3D生成风格迁移扩散模型多视角一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。