提出新方法分离视频的外观与运动,实现更灵活的视频生成控制。
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
- 用3D点云表示视频动态,通过多频编码区分细微运动。
- 在多种视频编辑任务中表现优于现有方法,支持相机与物体编辑。
- 适合需要精准控制视频生成的开发者与研究人员。
视频生成中的有效且可泛化的控制仍是重大挑战。现有方法多依赖模糊或特定任务的信号,我们主张通过“外观”与“运动”的基础解耦,提供更鲁棒、可扩展的路径。为此提出FlexAM,一种基于新型3D控制信号的统一框架。该信号将视频动态表示为点云,引入三项关键改进:多频位置编码以区分细粒度运动、深度感知编码,以及可灵活调节精度与泛化能力的控制信号。该表示使FlexAM能有效解耦外观与运动,支持包括图像到视频、视频到视频编辑、相机控制及空间物体编辑在内的多种任务。大量实验表明,FlexAM在所有评估任务中均取得更优性能。
原文摘要 · Abstract (English)
Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more robust and scalable pathway. We propose FlexAM, a unified framework built upon a novel 3D control signal. This signal represents video dynamics as a point cloud, introducing three key enhancements: multi-frequency positional encoding to distinguish fine-grained motion, depth-aware encoding, and a flexible control signal for balancing precision and generalization. This representation allows FlexAM to effectively disentangle appearance and motion, enabling a wide range of tasks including I2V/V2V editing, camera control, and spatial object editing. Extensive experiments demonstrate that FlexAM achieves superior performance across all evaluated tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。