让视频生成同时懂外观和运动,提升动态真实感。
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

- 用联合外观-运动表征替代传统像素重建目标。
- 在推理时用自生成运动信号动态引导,提升动作连贯性。
- 可适配任意视频模型,无需改数据或调规模,适合追求自然动态的开发者。
尽管生成式视频模型近期取得显著进展,但仍难以捕捉真实世界的运动、动态和物理规律。我们发现,这一局限源于传统的像素重建目标,它使模型更关注外观保真度而牺牲了运动连贯性。为此,我们提出VideoJAM,一种新框架,通过鼓励模型学习联合外观-运动表征来注入有效的运动先验。VideoJAM由两个互补模块组成:训练时扩展目标,从单一表征中预测生成像素及其对应运动;推理时引入Inner-Guidance机制,利用模型自身不断演化的运动预测作为动态引导信号,推动生成更连贯的运动。该框架可应用于任意视频模型,仅需最小改动,无需修改训练数据或调整模型规模。VideoJAM在运动连贯性上达到当前最优性能,超越多个高度竞争的专有模型,同时提升了生成结果的视觉质量。结果表明,外观与运动可互补,有效整合后能同时增强视频的视觉质量和动态一致性。
原文摘要 · Abstract (English)
Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。