arXiv:2607.23023cs.CV2026-07

实现交互式虚拟人实时音视频生成,支持长时间对话不跑偏。

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

论文配图:OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
图 1 · 摘自论文原文
  • 用进度控制器动态调节生成进度,支持流畅响应与倾听切换。
  • 多参考条件模块保持长时间音视频身份一致,避免形象漂移。
  • 适合开发虚拟助手、游戏角色等需要持续互动的场景。

基于扩散模型的生成技术已实现音频驱动的虚拟人实时生成和统一音视频合成,为交互式虚拟人系统奠定基础。然而,将统一音视频合成扩展到实时交互流仍面临挑战:生成时长未知,且长期生成中身份易漂移。为此,我们提出OmniMate,一个面向开放性实时交互的统一音视频生成框架。OmniMate实时联合生成视觉内容、语音与音效,支持自然沉浸的多轮交互。为实现自适应响应推进,引入生成进度控制器(GPC),显式建模每个流式片段的生成进度,使模型可根据目标进度完成响应,并实现执行与倾听状态间的无缝切换。为保持长期跨模态身份一致性,提出多参考条件模块(MRCM),利用多个参考图像和参考语音段提供持续的视觉与说话人身份线索。在交互导向的VerseBench适配数据集上的大量实验表明,OmniMate实现了高质量、低延迟的流式生成,同时保持强长期音视频一致性。结果进一步显示,OmniMate可在长时间多轮对话中支持真实、连贯且响应迅速的交互体验。

原文摘要 · Abstract (English)

Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending unified audio-visual synthesis to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and the generated identity may drift over long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and sound effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.

虚拟人音视频生成实时交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。