用少量数据实现多人互动视频生成,支持任意扩展身份数。
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- 通过身份感知注意力机制,逐对处理音视频流以扩展可驱动身份数。
- 仅需单人视频训练,用少量真人多人群体片段优化互动自然度。
- 提出新评估指标与数据集,专门衡量生成视频的自然性与互动性。
近期多人视频生成开始受到关注。尽管已有少数工作探索音频驱动的多人对话视频生成,但普遍面临多样多人数据采集成本高、多身份协同驱动困难等问题。为此,我们提出AnyTalker,一种可扩展的多流处理框架。具体而言,将扩散变换器的注意力模块扩展为新型身份感知注意力机制,迭代处理身份-音频对,实现可任意扩展的驱动身份数。此外,训练多人生成模型需要大量多人数据,而我们的训练流程仅依赖单人视频学习多人说话模式,并通过少量真实多人片段精修互动效果。同时,我们构建了针对性评估指标与数据集,用于评测生成视频的自然度与互动性。大量实验表明,AnyTalker在唇部同步、视觉质量及自然互动方面表现优异,在数据成本与身份可扩展性间取得良好平衡。
原文摘要 · Abstract (English)
Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often face challenges due to the high costs of diverse multi-person data collection and the difficulty of driving multiple identities with coherent interactivity. To address these challenges, we propose AnyTalker, a multi-person generation framework that features an extensible multi-stream processing architecture. Specifically, we extend Diffusion Transformer's attention block with a novel identity-aware attention mechanism that iteratively processes identity-audio pairs, allowing arbitrary scaling of drivable identities. Besides, training multi-person generative models demands massive multi-person data. Our proposed training pipeline depends solely on single-person videos to learn multi-person speaking patterns and refines interactivity with only a few real multi-person clips. Furthermore, we contribute a targeted metric and dataset designed to evaluate the naturalness and interactivity of the generated multi-person videos. Extensive experiments demonstrate that AnyTalker achieves remarkable lip synchronization, visual quality, and natural interactivity, striking a favorable balance between data costs and identity scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。