arXiv:2509.05285cs.GRcs.CV2025-09被引 1

用文本控制3D场景风格迁移,实现多视角一致且区域可控的高质量风格化。

Improved 3D Scene Stylization via Text-Guided Generative Image Editing with Region-Based Control

  • 基于单参考注意力机制,提升多视角间风格一致性。
  • 引入多深度图参考,增强不同视角生成图像的一致性。
  • 通过分割掩码实现区域级风格迁移,支持混合风格应用。

近期基于文本驱动的3D场景编辑与风格化方法借助2D生成模型的强大能力,取得了良好效果。然而,如何同时保证高质量风格化与多视角一致性仍是挑战,尤其在不同区域或物体间实现语义对应的一致风格迁移更具难度。为此,本文提出一种新方法:通过重训练初始3D表示,利用源视角的风格化多视图2D图像实现3D风格化。为确保风格与视角一致性,我们扩展了风格对齐的深度条件视图生成框架,将全共享注意力替换为单参考注意力共享机制,有效对齐不同视角的风格。此外,受3D修复方法启发,采用多个深度图组成的网格作为单图像参考,进一步强化风格化图像间的视图一致性。最后,提出多区域重要性加权切片沃尔什距离损失(Multi-Region Importance-Weighted Sliced Wasserstein Distance Loss),结合现成模型的分割掩码,实现不同图像区域的风格独立控制。实验表明,该方法显著提升了文本驱动3D风格化的质量与真实性,支持跨区域风格混合。项目页面:https://haruolabs.github.io/improved-gs-style-page/

原文摘要 · Abstract (English)

Recent advances in text-driven 3D scene editing and stylization, which leverage the powerful capabilities of 2D generative models, have demonstrated promising outcomes. However, challenges remain in ensuring high-quality stylization and view consistency simultaneously. Moreover, applying style consistently to different regions or objects in the scene with semantic correspondence is a challenging task. To address these limitations, we introduce techniques that enhance the quality of 3D stylization while maintaining view consistency and providing optional region-controlled style transfer. Our method achieves stylization by re-training an initial 3D representation using stylized multi-view 2D images of the source views. Therefore, ensuring both style consistency and view consistency of stylized multi-view images is crucial. We achieve this by extending the style-aligned depth-conditioned view generation framework, replacing the fully shared attention mechanism with a single reference-based attention-sharing mechanism, which effectively aligns style across different viewpoints. Additionally, inspired by recent 3D inpainting methods, we utilize a grid of multiple depth maps as a single-image reference to further strengthen view consistency among stylized images. Finally, we propose Multi-Region Importance-Weighted Sliced Wasserstein Distance Loss, allowing styles to be applied to distinct image regions using segmentation masks from off-the-shelf models. We demonstrate that this optional feature enhances the faithfulness of style transfer and enables the mixing of different styles across distinct regions of the scene. Experimental evaluations, both qualitative and quantitative, demonstrate that our pipeline effectively improves the results of text-driven 3D stylization. Project Page: https://haruolabs.github.io/improved-gs-style-page/

3D风格化文本控制多视角一致区域控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。