用2D扩散模型生成可多视角观看的3D视觉幻觉,支持文本或图像输入。
Illusion3D: 3D Multiview Illusion with 2D Diffusion Priors
- 基于预训练2D扩散模型优化神经3D结构的纹理与几何
- 多视角观看时呈现不同视觉解释,实现动态幻觉效果
- 适合艺术创作与交互设计,支持复杂3D形式生成
自动生成多视角视觉幻觉是一项引人注目的挑战,即单一视觉内容在不同视角下呈现不同解读。传统方法如阴影艺术和线框艺术虽能生成有趣3D幻觉,但仅限于简单视觉输出(如图底分离或线条绘制),限制了其艺术表现力和实际应用性。近期基于扩散模型的幻觉生成方法可生成更复杂设计,但局限于2D图像。本文提出一种简单而有效的方法,基于用户提供的文本提示或图像生成3D多视角幻觉。该方法利用预训练的文生图扩散模型,通过可微渲染优化神经3D表示的纹理与几何结构。从多个角度观察时,可产生不同视觉解读。我们开发了多种技术以提升生成幻觉的质量。通过大量实验验证了该方法的有效性,并展示了多样化3D形态的幻觉生成结果。
原文摘要 · Abstract (English)
Automatically generating multiview illusions is a compelling challenge, where a single piece of visual content offers distinct interpretations from different viewing perspectives. Traditional methods, such as shadow art and wire art, create interesting 3D illusions but are limited to simple visual outputs (i.e., figure-ground or line drawing), restricting their artistic expressiveness and practical versatility. Recent diffusion-based illusion generation methods can generate more intricate designs but are confined to 2D images. In this work, we present a simple yet effective approach for creating 3D multiview illusions based on user-provided text prompts or images. Our method leverages a pre-trained text-to-image diffusion model to optimize the textures and geometry of neural 3D representations through differentiable rendering. When viewed from multiple angles, this produces different interpretations. We develop several techniques to improve the quality of the generated 3D multiview illusions. We demonstrate the effectiveness of our approach through extensive experiments and showcase illusion generation with diverse 3D forms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。