arXiv:2507.03905cs.CV2025-07AAAI被引 31

13亿参数统一多模态多任务人体动画,效率远超大模型。

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

  • 用多任务掩码和反直觉分配策略,一模型跑多种动画任务。
  • 1.3B参数实现媲美大模型的动画质量,推理更快更省资源。
  • 适合需要高效多任务动画生成的研究与工业应用。

近期的人体动画研究通常依赖大规模视频模型,虽表现生动但存在推理慢、计算开销大的问题。传统方法对每项任务使用独立模型,多任务场景成本高且效率低。为此,我们提出EchoMimicV3,一个统一多任务、多模态人体动画的高效框架。核心设计包括:多任务混合(Soup-of-Tasks)机制,通过多任务掩码与反直觉任务分配策略,在不增加模型数量的前提下实现多任务收益;多模态混合(Soup-of-Modals)机制,引入耦合-解耦多模态交叉注意力模块注入多模态条件,并结合时序感知的多模态动态分配机制调节多模态融合。此外,我们提出负向直接偏好优化、时序感知负向无分类器引导(CFG)及长视频CFG,保障训练与推理稳定。大量实验表明,仅1.3亿参数的EchoMimicV3在定量与定性评估中均达到竞争力水平。

原文摘要 · Abstract (English)

Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for each animation task, increasing costs in multi-task scenarios and worsening the dilemma. To address these limitations, we introduce EchoMimicV3, an efficient framework that unifies multi-task and multi-modal human animation. At the core of EchoMimicV3 lies a threefold design: a Soup-of-Tasks paradigm, a Soup-of-Modals paradigm, and a novel training and inference strategy. The Soup-of-Tasks leverages multi-task mask inputs and a counter-intuitive task allocation strategy to achieve multi-task gains without multi-model pains. Meanwhile, the Soup-of-Modals introduces a Coupled-Decoupled Multi-Modal Cross Attention module to inject multi-modal conditions, complemented by a Multi-Modal Timestep Phase-aware Dynamical Allocation mechanism to modulate multi-modal mixtures. Besides, we propose Negative Direct Preference Optimization, Phase-aware Negative Classifier-Free Guidance (CFG), and Long Video CFG, which ensure stable training and inference. Extensive experiments and analyses demonstrate that EchoMimicV3, with a minimal model size of 1.3 billion parameters, achieves competitive performance in both quantitative and qualitative evaluations.

人体动画多任务多模态高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。