让用户直观控制视频中镜头与物体运动,实现电影级运镜生成。
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation

- 通过场景感知的控制方式联合调节镜头与物体运动。
- 无需3D训练数据即可实现3D-aware运动控制。
- 适合影视创作、视频编辑等需要精细运动设计的场景。
本文提出MotionCanvas,一种在图像到视频生成框架下实现电影级镜头设计的方法。镜头设计是影视创作的关键,涉及对镜头运动和场景内物体运动的精确规划。现有图像到视频生成系统在实现直观镜头设计时面临两大挑战:一是如何有效捕捉用户对运动意图的表达,需同时指定镜头运动与物体运动;二是如何表示可被视频扩散模型有效利用的运动信息。为此,MotionCanvas将用户驱动的控制融入图像到视频生成模型,实现对场景中物体与镜头运动的场景感知式控制。结合传统计算机图形学与现代视频生成技术,该方法在不依赖昂贵3D训练数据的前提下,实现了3D感知的运动控制。用户可直观描绘场景空间中的运动意图,并将其转化为时空运动条件信号输入视频扩散模型。我们在多种真实图像内容与镜头设计场景中验证了该方法的有效性,展现了其在数字内容创作流程中的增强潜力,以及在图像与视频编辑应用中的广泛适应性。
原文摘要 · Abstract (English)
This paper presents a method that allows users to design cinematic video shots in the context of image-to-video generation. Shot design, a critical aspect of filmmaking, involves meticulously planning both camera movements and object motions in a scene. However, enabling intuitive shot design in modern image-to-video generation systems presents two main challenges: first, effectively capturing user intentions on the motion design, where both camera movements and scene-space object motions must be specified jointly; and second, representing motion information that can be effectively utilized by a video diffusion model to synthesize the image animations. To address these challenges, we introduce MotionCanvas, a method that integrates user-driven controls into image-to-video (I2V) generation models, allowing users to control both object and camera motions in a scene-aware manner. By connecting insights from classical computer graphics and contemporary video generation techniques, we demonstrate the ability to achieve 3D-aware motion control in I2V synthesis without requiring costly 3D-related training data. MotionCanvas enables users to intuitively depict scene-space motion intentions, and translates them into spatiotemporal motion-conditioning signals for video diffusion models. We demonstrate the effectiveness of our method on a wide range of real-world image content and shot-design scenarios, highlighting its potential to enhance the creative workflows in digital content creation and adapt to various image and video editing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。