arXiv:2512.03683cs.CV2025-12

一键实现高保真3D风格化,无需逐项优化。

GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces

  • 通过解耦潜在空间,分离几何与外观信息进行快速编辑
  • 推理时即时生成,多视角一致且保留原始结构
  • 适合游戏开发与数字艺术的大规模风格化需求

3D风格化在游戏开发、虚拟现实和数字艺术中至关重要,需高效可扩展的方法以支持快速、高保真操作。现有文本驱动3D风格化方法通常基于2D图像编辑器迁移,需耗时的逐资产优化,且受当前文本到图像模型限制,导致多视角不一致,难以用于大规模生产。本文提出GaussianBlender,首个前馈式文本驱动3D风格化框架,可在推理时即时完成编辑。该方法从空间分组的3D Gaussians中学习结构化、解耦的潜在空间,控制几何与外观间的信息共享;再利用潜在扩散模型对这些表示施加文本条件编辑。全面评估表明,GaussianBlender不仅实现即时、高保真、几何保持、多视角一致的风格化,还超越需实例级测试时优化的方法,真正实现规模化、普惠化的3D风格化。

原文摘要 · Abstract (English)

3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylization methods typically distill from 2D image editors, requiring time-intensive per-asset optimization and exhibiting multi-view inconsistency due to the limitations of current text-to-image models, which makes them impractical for large-scale production. In this paper, we introduce GaussianBlender, a pioneering feed-forward framework for text-driven 3D stylization that performs edits instantly at inference. Our method learns structured, disentangled latent spaces with controlled information sharing for geometry and appearance from spatially-grouped 3D Gaussians. A latent diffusion model then applies text-conditioned edits on these learned representations. Comprehensive evaluations show that GaussianBlender not only delivers instant, high-fidelity, geometry-preserving, multi-view consistent stylization, but also surpasses methods that require per-instance test-time optimization - unlocking practical, democratized 3D stylization at scale.

3D风格化扩散模型即时编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。