让对话模型生成带表情动作的视频回复,实现声音与画面同步
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

- 通过视觉思维计划统一控制文本、语音和视频生成
- 四步推理下端到端实时率1.293,400×720分辨率可流畅运行
- 适合需要真实感虚拟人交互的研究者与开发者
全能模态对话模型能理解多模态输入并生成语音回复,但其输出缺乏视觉表现力。我们提出Ex-Omni-2D,一个能生成包含文本、个性化语音和参考条件视频的协调响应框架。给定多模态查询、参考图像和音频,模型首先预测描述场景、情绪和动作的结构化视觉思维计划(VTP),随后生成响应文本和原生多码本语音单元。这些语音单元构成共享声学-时间接口:既可解码为语音,又能在线对齐视频帧。该接口使响应与虚拟形象路径可从异构的语音、对话和虚拟视频数据中联合学习,无需大规模查询-文本-语音-视频标注。主干采用全序列视频生成器作为教师模型;为提升效率,进一步蒸馏为几步块因果的流式学生模型,其前缀流机制在连续分块间传递干净潜在表示,减少累积延迟误差。四步推理下,四卡全流程实现端到端实时率1.293,分辨率400×720/720×400,达到实用的质量-效率平衡点。
原文摘要 · Abstract (English)
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured \textit{Visual Thought Plan} (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal \emph{Streaming Student} whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at $400\times720$/$720\times400$, providing a practical quality--efficiency operating point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。