首个支持多模态风格迁移的通用框架,直接作用于高斯点云。
CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian Splatting
- 在高斯点云上直接进行文本或图像引导的风格迁移。
- 支持2D/3D/4D内容,保持时序一致性且不增加模型体积。
- 无需重训练,可作为插件集成到现有高斯渲染流程中。
高斯点阵(Gaussian Splatting, GS)最近成为从2D图像高效重建3D场景的代表方法,并已扩展至图像、视频及动态4D内容。然而,将风格迁移应用于基于GS的表示,尤其是超越简单色彩调整的任务,仍面临挑战。本文提出CLIPGaussian,首个统一的多模态风格迁移框架,支持文本与图像引导的风格化,适用于2D图像、视频、3D物体和4D场景。该方法直接在高斯原始数据上操作,可作为即插即用模块集成到现有GS管线中,无需大型生成模型或从头训练。通过联合优化颜色与几何,实现3D/4D场景中的风格一致性,并在视频中保持时序连贯性,同时维持模型大小不变。实验表明,该方法在各类任务中均表现出更优的风格保真度与一致性,验证了其作为通用高效多模态风格迁移方案的潜力。
原文摘要 · Abstract (English)
Gaussian Splatting (GS) has recently emerged as an efficient representation for rendering 3D scenes from 2D images and has been extended to images, videos, and dynamic 4D content. However, applying style transfer to GS-based representations, especially beyond simple color changes, remains challenging. In this work, we introduce CLIPGaussian, the first unified style transfer framework that supports text- and image-guided stylization across multiple modalities: 2D images, videos, 3D objects, and 4D scenes. Our method operates directly on Gaussian primitives and integrates into existing GS pipelines as a plug-in module, without requiring large generative models or retraining from scratch. The CLIPGaussian approach enables joint optimization of color and geometry in 3D and 4D settings, and achieves temporal coherence in videos, while preserving the model size. We demonstrate superior style fidelity and consistency across all tasks, validating CLIPGaussian as a universal and efficient solution for multimodal style transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。