arXiv:2507.05963cs.CV2025-07被引 12

Tora2实现多角色外观与动作同步定制,提升视频生成精度与可控性。

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

  • 解耦个性化提取器生成多实体细粒度特征嵌入
  • 引入门控自注意力机制减少多模态条件错位
  • 对比损失联合优化动作与身份一致性,适合需要精细控制的视频生成场景

近期基于扩散变换器的运动引导视频生成模型(如Tora)取得了显著进展。本文提出Tora2,作为Tora的增强版本,通过多项设计改进,显著提升了在外观和运动定制方面的能力。具体而言,我们引入解耦的个性化提取器,为多个开放集实体生成全面的个性化嵌入,相比之前方法更好地保留了细粒度视觉细节。在此基础上,设计门控自注意力机制,整合每个实体的轨迹、文本描述与视觉信息,有效降低训练中多模态条件的错位问题。此外,提出一种对比损失,通过显式映射运动嵌入与个性化嵌入,联合优化轨迹动态与实体一致性。据我们所知,Tora2是首个实现多实体外观与运动同步定制的视频生成方法。实验表明,Tora2在保持与先进定制方法相当性能的同时,具备更强的动作控制能力,标志着多条件视频生成的重要进展。

原文摘要 · Abstract (English)

Recent advances in diffusion transformer models for motion-guided video generation, such as Tora, have shown significant progress. In this paper, we present Tora2, an enhanced version of Tora, which introduces several design improvements to expand its capabilities in both appearance and motion customization. Specifically, we introduce a decoupled personalization extractor that generates comprehensive personalization embeddings for multiple open-set entities, better preserving fine-grained visual details compared to previous methods. Building on this, we design a gated self-attention mechanism to integrate trajectory, textual description, and visual information for each entity. This innovation significantly reduces misalignment in multimodal conditioning during training. Moreover, we introduce a contrastive loss that jointly optimizes trajectory dynamics and entity consistency through explicit mapping between motion and personalization embeddings. Tora2 is, to our best knowledge, the first method to achieve simultaneous multi-entity customization of appearance and motion for video generation. Experimental results demonstrate that Tora2 achieves competitive performance with state-of-the-art customization methods while providing advanced motion control capabilities, which marks a critical advancement in multi-condition video generation. Project page: https://ali-videoai.github.io/Tora2_page/.

视频生成扩散模型多实体控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。