用空间分解建模实现角色视频的可控生成
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
- 将2D视频分解为人物、场景、遮挡三部分,基于3D深度分层编码
- 支持任意角色快速生成,可控制动作与场景交互
- 适合需要灵活控制角色动作和环境互动的应用场景
角色视频合成旨在生成可动画化角色在逼真场景中的真实视频。作为计算机视觉与图形学领域的基础问题,传统3D方法通常需要多视角采集进行逐例训练,严重限制了对任意角色的快速建模。近期2D方法通过预训练扩散模型突破此限制,但在动作泛化性和场景交互方面表现不佳。为此,本文提出MIMO框架,可在统一架构下实现对角色、动作和场景的可控生成,同时具备对任意角色的高扩展性、对新3D动作的强泛化能力以及对交互式真实场景的适用性。核心思想是将2D视频编码为紧凑的空间代码,考虑视频发生的固有3D特性。具体地,利用单目深度估计将2D帧像素升维至3D,基于3D深度将视频片段分层解构为三类空间成分:主体人物、底层场景和浮动遮挡。这些成分进一步被编码为规范身份码、结构化动作码和完整场景码,作为合成过程的控制信号。空间分解建模设计实现了灵活用户控制、复杂动作表达及对场景交互的3D感知合成。实验结果验证了该方法的有效性与鲁棒性。
原文摘要 · Abstract (English)
Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。