arXiv:2504.21836cs.CV2025-04International Conf…被引 8

用大模型注意力机制实现3D物体风格迁移,无需训练即可高效保持多视角一致性。

3D Stylization via Large Reconstruction Model

  • 利用3D生成模型中特定注意力层提取风格特征
  • 注入参考图像特征后实现高质量3D风格迁移
  • 无需训练或优化,适合快速原型设计与内容创作

随着文本或图像引导的3D生成技术日益成熟,用户对生成过程的控制力需求提升,其中外观风格化尤为关键。给定参考图像,需将生成3D资产的外观适配至参考图的视觉风格,同时保证多视角间的一致性。受2D风格迁移中大型图像生成模型注意力机制启发,本文探究大型重建模型是否具备类似能力。研究发现,此类模型中的部分注意力模块可捕捉特定外观特征。通过向这些模块注入视觉风格图像的特征,我们提出一种无需训练或测试时优化的简单而有效的3D外观风格化方法。定量与定性评估表明,该方法在3D外观风格化任务中表现优异,显著提升效率并保持高质量输出。

原文摘要 · Abstract (English)

With the growing success of text or image guided 3D generators, users demand more control over the generation process, appearance stylization being one of them. Given a reference image, this requires adapting the appearance of a generated 3D asset to reflect the visual style of the reference while maintaining visual consistency from multiple viewpoints. To tackle this problem, we draw inspiration from the success of 2D stylization methods that leverage the attention mechanisms in large image generation models to capture and transfer visual style. In particular, we probe if large reconstruction models, commonly used in the context of 3D generation, has a similar capability. We discover that the certain attention blocks in these models capture the appearance specific features. By injecting features from a visual style image to such blocks, we develop a simple yet effective 3D appearance stylization method. Our method does not require training or test time optimization. Through both quantitative and qualitative evaluations, we demonstrate that our approach achieves superior results in terms of 3D appearance stylization, significantly improving efficiency while maintaining high-quality visual outcomes.

3D生成风格迁移注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。