让3D场景一键换风格,支持文字或图片控制
AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting
- 用文本或参考图直接控制3D场景风格
- 无需姿态信息,零样本适配新风格
- 可无缝接入现有3D重建模型,适合快速创作
随着对快速、可扩展3D资产生成的需求增长,前馈式3D重建方法受到关注,其中3D高斯点云(3DGS)已成为有效的场景表示方式。尽管近期方法已实现无姿态图像集合的重建,但将风格化或外观控制集成到此类流程中仍研究不足。现有方法多依赖图像条件,限制了可控性与灵活性。本文提出AnyStyle,一种前馈式3D重建与风格化框架,通过多模态条件实现无姿态、零样本风格化。该方法支持文本和视觉风格输入,用户可用自然语言描述或参考图像控制场景外观。我们设计模块化风格化架构,仅需少量结构修改,即可融入现有前馈3D重建主干网络。实验表明,AnyStyle在风格可控性上优于先前前馈风格化方法,同时保持高质量几何重建。用户研究进一步验证,其风格化质量优于现有最先进方法。
原文摘要 · Abstract (English)
The growing demand for rapid and scalable 3D asset creation has driven interest in feed-forward 3D reconstruction methods, with 3D Gaussian Splatting (3DGS) emerging as an effective scene representation. While recent approaches have demonstrated pose-free reconstruction from unposed image collections, integrating stylization or appearance control into such pipelines remains underexplored. Existing attempts largely rely on image-based conditioning, which limits both controllability and flexibility. In this work, we introduce AnyStyle, a feed-forward 3D reconstruction and stylization framework that enables pose-free, zero-shot stylization through multimodal conditioning. Our method supports both textual and visual style inputs, allowing users to control the scene appearance using natural language descriptions or reference images. We propose a modular stylization architecture that requires only minimal architectural modifications and can be integrated into existing feed-forward 3D reconstruction backbones. Experiments demonstrate that AnyStyle improves style controllability over prior feed-forward stylization methods while preserving high-quality geometric reconstruction. A user study further confirms that AnyStyle achieves superior stylization quality compared to an existing state-of-the-art approach. Repository: https://github.com/joaxkal/AnyStyle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。