arXiv:2604.08760cs.CV2026-04

用图像控制文本生成3D模型,提升细节和风格一致性。

SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation

  • 分两阶段生成:先文本转3D高斯点云,再用参考图迁移风格。
  • 新设计的损失函数有效保留全局与局部纹理,减少几何与外观冲突。
  • 在几何精度和风格还原上均优于现有方法,适合需要精细控制的场景。

近期文本到3D物体生成进展通过结合2D扩散模型与可微3D表示,实现了从文本输入生成细节丰富的几何结构。然而,由于文本模态的局限性,现有方法常面临可控性差和纹理模糊的问题。为此,我们提出SIC3D,一种基于3D高斯点云(3DGS)的图像条件化文本到3D生成流程。SIC3D包含两个阶段:第一阶段使用文本到3DGS生成模型从文本生成3D物体内容;第二阶段将参考图像的风格迁移到3DGS中。在风格迁移阶段,我们引入一种新型变分风格化得分蒸馏(VSSD)损失,有效捕捉全局与局部纹理模式,同时缓解几何与外观之间的冲突。此外,采用缩放正则化防止伪影出现,并保留风格图像中的图案信息。大量实验表明,SIC3D在几何保真度和风格一致性的定量与定性评估中均优于现有方法。

原文摘要 · Abstract (English)

Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differentiable 3D representations. However, the approaches often suffer from limited controllability and texture ambiguity due to the limitation of the text modality. To address this, we present SIC3D, a controllable image-conditioned text-to-3D generation pipeline with 3D Gaussian Splatting (3DGS). There are two stages in SIC3D. The first stage generates the 3D object content from text with a text-to-3DGS generation model. The second stage transfers style from a reference image to the 3DGS. Within this stylization stage, we introduce a novel Variational Stylized Score Distillation (VSSD) loss to effectively capture both global and local texture patterns while mitigating conflicts between geometry and appearance. A scaling regularization is further applied to prevent the emergence of artifacts and preserve the pattern from the style image. Extensive experiments demonstrate that SIC3D enhances geometric fidelity and style adherence, outperforming prior approaches in both qualitative and quantitative evaluations.

3D生成风格迁移高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。