arXiv:2501.12935cs.CV2025-01被引 8

用生成模型实现单图3D物体的精准编辑与逼真动态效果

3D Object Manipulation in a Single Image using Generative Models

  • 将2D图像转为3D,结合扩散模型实现几何级精细控制
  • 自研纹理优化模块使渲染细节与原图风格高度一致
  • 自动匹配背景光照,生成符合人眼感知的真实阴影

图像中的物体操作不仅需修改物体外观,还需赋予其运动能力。以往方法难以同时实现静态编辑与动态生成,且在物体外观和场景光照真实性上表现不佳。本文提出 OMG3D 框架,将精确几何控制与扩散模型的生成能力融合,在视觉表现上取得显著提升。该框架首先将2D物体转换为3D,支持用户引导的修改与逼真运动。为提升纹理真实感,提出 CustomRefiner 模块,预训练定制化扩散模型,使3D粗略模型的渲染细节与原始图像对齐并进一步优化。此外,引入 IllumiCombiner 光照处理模块,估计并校正背景光照,使其符合人类视觉感知,从而产生更真实的阴影效果。大量实验表明,本方法在静态与动态场景中均表现出色,所有步骤均可在单张 NVIDIA 3090 上完成。

原文摘要 · Abstract (English)

Object manipulation in images aims to not only edit the object's presentation but also gift objects with motion. Previous methods encountered challenges in concurrently handling static editing and dynamic generation, while also struggling to achieve fidelity in object appearance and scene lighting. In this work, we introduce \textbf{OMG3D}, a novel framework that integrates the precise geometric control with the generative power of diffusion models, thus achieving significant enhancements in visual performance. Our framework first converts 2D objects into 3D, enabling user-directed modifications and lifelike motions at the geometric level. To address texture realism, we propose CustomRefiner, a texture refinement module that pre-train a customized diffusion model, aligning the details and style of coarse renderings of 3D rough model with the original image, further refine the texture. Additionally, we introduce IllumiCombiner, a lighting processing module that estimates and corrects background lighting to match human visual perception, resulting in more realistic shadow effects. Extensive experiments demonstrate the outstanding visual performance of our approach in both static and dynamic scenarios. Remarkably, all these steps can be done using one NVIDIA 3090. Project page is at https://whalesong-zrs.github.io/OMG3D-projectpage/

3D生成图像编辑扩散模型光照建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。