让虚拟人动得有情感有逻辑,不只是模仿动作。
OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- 用大模型生成语义引导文本,让动作更懂情绪和场景
- 融合音视频文三模态,动作自然且与提示语一致
- 适合做有性格的虚拟角色,尤其多角色复杂场景
现有视频虚拟人模型虽能生成流畅动作,但仅依赖音频节奏等低层信号,缺乏对情感、意图和上下文的深层理解。为填补这一差距,本文提出OmniHuman-1.5框架,通过多模态大语言模型生成结构化语义提示,指导运动生成器实现语义连贯的动作表达。同时引入专用的多模态DiT架构与伪末帧设计,有效融合音频、图像与文本信息,缓解跨模态冲突。实验表明,该模型在唇同步准确率、视频质量、动作自然度及文本语义一致性等多项指标上均达领先水平。方法还可扩展至多人及非人类主体的复杂场景,表现出强泛化能力。
原文摘要 · Abstract (English)
Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm, lacking a deeper semantic understanding of emotion, intent, or context. To bridge this gap, \textbf{we propose a framework designed to generate character animations that are not only physically plausible but also semantically coherent and expressive.} Our model, \textbf{OmniHuman-1.5}, is built upon two key technical contributions. First, we leverage Multimodal Large Language Models to synthesize a structured textual representation of conditions that provides high-level semantic guidance. This guidance steers our motion generator beyond simplistic rhythmic synchronization, enabling the production of actions that are contextually and emotionally resonant. Second, to ensure the effective fusion of these multimodal inputs and mitigate inter-modality conflicts, we introduce a specialized Multimodal DiT architecture with a novel Pseudo Last Frame design. The synergy of these components allows our model to accurately interpret the joint semantics of audio, images, and text, thereby generating motions that are deeply coherent with the character, scene, and linguistic content. Extensive experiments demonstrate that our model achieves leading performance across a comprehensive set of metrics, including lip-sync accuracy, video quality, motion naturalness and semantic consistency with textual prompts. Furthermore, our approach shows remarkable extensibility to complex scenarios, such as those involving multi-person and non-human subjects. Homepage: \href{https://omnihuman-lab.github.io/v1_5/}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。