arXiv:2508.03034cs.CV2025-08被引 5

MoCA通过混合交叉注意力实现高保真人脸视频生成

MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention

  • 引入混合交叉注意力机制,提升跨帧身份一致性
  • 在CelebIPVid数据集上身份相似度超越现有方法5%以上
  • 适合需要精确人脸保持的视频生成场景

尽管基于扩散模型的文本到视频生成取得进展,但保持身份一致仍具挑战。现有方法常无法捕捉细微面部动态或维持时间上的一致性。为此,我们提出MoCA,一种基于扩散Transformer(DiT)骨干网络的视频扩散模型,采用受专家混合范式启发的混合交叉注意力机制。框架在每个DiT模块中嵌入MoCA层,通过分层时序池化捕捉不同时间尺度的身份特征,并利用时序感知交叉注意力专家动态建模时空关系。此外,引入潜在视频感知损失以增强跨帧身份一致性和细节表现。为训练该模型,我们收集了包含1000名不同族裔个体的10,000段高清视频的CelebIPVid数据集,促进跨族裔泛化能力。在CelebIPVid上的大量实验表明,MoCA在身份相似度上比现有T2V方法提升超过5%。

原文摘要 · Abstract (English)

Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnicity generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across Face similarity.

视频生成扩散模型身份保持注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。