arXiv:2604.03773cs.CV2026-04

用文字或图片做风格参考,实时生成3D风格化视图。

M2StyleGS: Multi-Modality 3D Style Transfer with Gaussian Splatting

  • 结合CLIP多模态特征与3D高斯点云实现风格迁移。
  • 引入分割流对齐与观察损失,提升风格一致性达32.92%。
  • 适合虚拟现实、增强现实等需灵活输入的场景使用。

传统3D风格迁移依赖固定参考图像,难以满足虚拟现实等场景中用户对文本描述或多样图像输入的灵活需求。本文提出M2StyleGS,基于3D高斯点云(3DGS)构建3D表现形式,利用CLIP提取的多模态知识作为风格参考。通过细分流(subdivisive flow)实现精确特征对齐,强化文本-视觉联合特征向VGG风格特征的投影。引入观察损失以提升生成场景与参考风格的一致性,抑制损失则在解码过程中防止参考色彩信息偏移。实验表明,该方法可支持文本或图像为参考,生成一组风格增强的新视角,视觉质量更优,一致性相比此前工作最高提升32.92%。

原文摘要 · Abstract (English)

Conventional 3D style transfer methods rely on a fixed reference image to apply artistic patterns to 3D scenes. However, in practical applications such as virtual or augmented reality, users often prefer more flexible inputs, including textual descriptions and diverse imagery. In this work, we introduce a novel real-time styling technique M2StyleGS to generate a sequence of precisely color-mapped views. It utilizes 3D Gaussian Splatting (3DGS) as a 3D presentation and multi-modality knowledge refined by CLIP as a reference style. M2StyleGS resolves the abnormal transformation issue by employing a precise feature alignment, namely subdivisive flow, it strengthens the projection of the mapped CLIP text-visual combination feature to the VGG style feature. In addition, we introduce observation loss, which assists in the stylized scene better matching the reference style during the generation, and suppression loss, which suppresses the offset of reference color information throughout the decoding process. By integrating these approaches, M2StyleGS can employ text or images as references to generate a set of style-enhanced novel views. Our experiments show that M2StyleGS achieves better visual quality and surpasses the previous work by up to 32.92% in terms of consistency.

3D风格迁移多模态高斯点云实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。