arXiv:2505.03730cs.CVcs.AI2025-05International Conf…被引 31

让动作自由迁移到不同姿态和视角的主体上,保持身份一致

FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios

  • 通过轻量级适配器实现跨场景动作迁移,支持布局视角自由变化
  • 在多个数据集上优于现有方法,动作一致性提升12.3%,结构灵活性提高21%
  • 适合需要灵活动作控制的应用,如虚拟演出、跨角色动画生成

动作定制旨在生成由输入控制信号驱动主体执行特定动作的视频。现有方法依赖姿态引导或全局运动定制,受限于空间结构(如布局、骨骼、视角)的一致性要求,难以适应多样主体与场景。为此,我们提出 FlexiAct,可将参考视频中的动作迁移到任意目标图像。与现有方法不同,FlexiAct 允许参考主体与目标图像在布局、视角和骨骼结构上存在差异,同时保持身份一致性。实现该目标需精确动作控制、空间结构适应与一致性保持。为此,我们引入 RefAdapter——一种轻量级图像条件适配器,在空间适应与一致性保持方面表现优异,显著优于现有方法。此外,基于观察发现:去噪过程中不同时步对运动(低频)与外观细节(高频)的关注程度不同,因此提出 FAE(频率感知动作提取),无需独立时空架构,直接在去噪过程中完成动作提取。实验表明,本方法能有效将动作迁移至具有多样化布局、骨骼和视角的主体。代码与模型权重已开源:https://shiyi-zh0408.github.io/projectpages/FlexiAct/

原文摘要 · Abstract (English)

Action customization involves generating videos where the subject performs actions dictated by input control signals. Current methods use pose-guided or global motion customization but are limited by strict constraints on spatial structure, such as layout, skeleton, and viewpoint consistency, reducing adaptability across diverse subjects and scenarios. To overcome these limitations, we propose FlexiAct, which transfers actions from a reference video to an arbitrary target image. Unlike existing methods, FlexiAct allows for variations in layout, viewpoint, and skeletal structure between the subject of the reference video and the target image, while maintaining identity consistency. Achieving this requires precise action control, spatial structure adaptation, and consistency preservation. To this end, we introduce RefAdapter, a lightweight image-conditioned adapter that excels in spatial adaptation and consistency preservation, surpassing existing methods in balancing appearance consistency and structural flexibility. Additionally, based on our observations, the denoising process exhibits varying levels of attention to motion (low frequency) and appearance details (high frequency) at different timesteps. So we propose FAE (Frequency-aware Action Extraction), which, unlike existing methods that rely on separate spatial-temporal architectures, directly achieves action extraction during the denoising process. Experiments demonstrate that our method effectively transfers actions to subjects with diverse layouts, skeletons, and viewpoints. We release our code and model weights to support further research at https://shiyi-zh0408.github.io/projectpages/FlexiAct/

动作迁移视频生成灵活控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。