零样本控制3D资产风格,可精细调节纹理与几何形态。
StyleSculptor: Zero-Shot Style-Controllable 3D Asset Generation with Texture-Geometry Dual Guidance
- 通过跨3D注意力机制动态融合内容与风格图像特征
- 支持仅纹理、仅几何或两者结合的风格控制,生成质量高
- 无需训练,适合游戏与VR领域快速生成定制化3D模型
在视频游戏和虚拟现实等实际应用中,生成符合已有图像纹理与几何风格的3D资产具有重要意义。尽管文本或图像到3D生成已取得显著进展,但实现风格可控的3D生成仍面临挑战。本文提出StyleSculptor,一种无需训练的零样本方法,可基于内容图像和一个或多个风格图像生成风格引导的3D资产。其核心是风格解耦注意力(SD-Attn)模块,通过跨3D注意力机制建立内容与风格图像间的动态交互,实现稳定特征融合与有效风格引导。为缓解语义内容泄露,该模块引入基于3D特征块方差的风格解耦特征选择策略,分离风格与内容相关通道,实现在注意力框架内选择性注入特征。借助此机制,网络可动态计算仅纹理、仅几何或两者兼具的引导特征以控制3D生成过程。进一步提出的风格引导控制(SGC)机制支持纯几何或纯纹理风格化及可调强度控制。大量实验表明,StyleSculptor在生成高质量3D资产方面优于现有基线方法。
原文摘要 · Abstract (English)
Creating 3D assets that follow the texture and geometry style of existing ones is often desirable or even inevitable in practical applications like video gaming and virtual reality. While impressive progress has been made in generating 3D objects from text or images, creating style-controllable 3D assets remains a complex and challenging problem. In this work, we propose StyleSculptor, a novel training-free approach for generating style-guided 3D assets from a content image and one or more style images. Unlike previous works, StyleSculptor achieves style-guided 3D generation in a zero-shot manner, enabling fine-grained 3D style control that captures the texture, geometry, or both styles of user-provided style images. At the core of StyleSculptor is a novel Style Disentangled Attention (SD-Attn) module, which establishes a dynamic interaction between the input content image and style image for style-guided 3D asset generation via a cross-3D attention mechanism, enabling stable feature fusion and effective style-guided generation. To alleviate semantic content leakage, we also introduce a style-disentangled feature selection strategy within the SD-Attn module, which leverages the variance of 3D feature patches to disentangle style- and content-significant channels, allowing selective feature injection within the attention framework. With SD-Attn, the network can dynamically compute texture-, geometry-, or both-guided features to steer the 3D generation process. Built upon this, we further propose the Style Guided Control (SGC) mechanism, which enables exclusive geometry- or texture-only stylization, as well as adjustable style intensity control. Extensive experiments demonstrate that StyleSculptor outperforms existing baseline methods in producing high-fidelity 3D assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。