MoCha首次实现电影级对话角色动画生成,让AI角色能自然说话并互动。
MoCha: Towards Movie-Grade Talking Character Synthesis
- 用语音-视频窗口注意力对齐声音与动作,确保口型精准同步。
- 联合训练语音与文本标注数据,提升角色动作泛化能力。
- 支持多角色轮换对话,生成具有剧情连贯性的影视级对话场景。
近期视频生成技术虽在运动真实感上取得进展,但常忽视角色驱动的叙事任务,而这正是自动化影视与动画生成的关键。本文提出「对话角色」(Talking Characters)这一更贴近真实的任务,旨在直接从语音和文本生成完整人物的动画,超越传统仅关注面部区域的「说话头」模型。我们提出首个此类系统MoCha,通过语音-视频窗口注意力机制,精确对齐语音与视频帧的时序特征。针对大规模语音标注视频数据稀缺问题,采用联合训练策略,同时利用语音与文本标注数据,显著提升模型在多样角色动作上的泛化性能。此外,设计带角色标签的结构化提示模板,首次实现多角色轮换对话,使生成角色可进行上下文感知的互动,保持电影级叙事连贯性。大量定性与定量评估(包括人类偏好实验与基准对比)表明,MoCha在真实感、表现力、可控性和泛化性方面均达到新高度,树立了AI电影叙事生成的新标准。
原文摘要 · Abstract (English)
Recent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation. We introduce Talking Characters, a more realistic task to generate talking character animations directly from speech and text. Unlike talking head, Talking Characters aims at generating the full portrait of one or more characters beyond the facial region. In this paper, we propose MoCha, the first of its kind to generate talking characters. To ensure precise synchronization between video and speech, we propose a speech-video window attention mechanism that effectively aligns speech and video tokens. To address the scarcity of large-scale speech-labeled video datasets, we introduce a joint training strategy that leverages both speech-labeled and text-labeled video data, significantly improving generalization across diverse character actions. We also design structured prompt templates with character tags, enabling, for the first time, multi-character conversation with turn-based dialogue-allowing AI-generated characters to engage in context-aware conversations with cinematic coherence. Extensive qualitative and quantitative evaluations, including human preference studies and benchmark comparisons, demonstrate that MoCha sets a new standard for AI-generated cinematic storytelling, achieving superior realism, expressiveness, controllability and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。