用3D几何控制让静态图动起来,还能自由调整镜头和动作。
I2V3D: Controllable image-to-video generation with 3D guidance

- 分两阶段生成:先精修关键帧,再用双向引导插值视频帧。
- 支持任意起始点和长序列动画,生成视频质量高且连贯。
- 适合需要精准控制动画镜头、角色动作的创作者或开发者。
我们提出 I2V3D,一种新颖的框架,可将静态图像生成动态视频,并实现精确的3D控制。该方法结合计算机图形学的精度与生成式AI的视觉质量,通过3D几何引导实现对相机运动、物体旋转及角色动画的准确控制。为支持任意起始点和长序列动画,采用两阶段生成流程:1)3D-Guided Keyframe Generation,使用定制图像扩散模型优化渲染关键帧以保证一致性和质量;2)3D-Guided Video Interpolation,一种无需训练的方法,利用双向引导在关键帧间生成平滑高质量视频帧。实验表明,该框架能从单张输入图生成可控且高质量的动画。代码将公开发布。
原文摘要 · Abstract (English)
We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a computer graphics pipeline, enabling accurate control over elements such as camera movement, object rotation, and character animation, with the visual fidelity of generative AI to produce high-quality videos from coarsely rendered inputs. To support animations with any initial start point and extended sequences, we adopt a two-stage generation process guided by 3D geometry: 1) 3D-Guided Keyframe Generation, where a customized image diffusion model refines rendered keyframes to ensure consistency and quality, and 2) 3D-Guided Video Interpolation, a training-free approach that generates smooth, high-quality video frames between keyframes using bidirectional guidance. Experimental results highlight the effectiveness of our framework in producing controllable, high-quality animations from single input images by harmonizing 3D geometry with generative models. The code for our framework will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。