arXiv:2508.19320cs.CVcs.AI2025-08被引 20

实时生成可多模态交互的数字人视频,低延迟且控制精准

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

  • 基于自回归框架,融合语音、姿态和文本多模态输入
  • 支持20,000小时对话数据训练,实现低延迟流式生成
  • 64倍压缩编码降低推理负担,适合实时交互场景

近期,交互式数字人视频生成受到广泛关注并取得显著进展。然而,构建能实时响应多样输入信号的实际系统仍具挑战,现有方法常面临计算开销大和控制能力有限的问题。本文提出一种自回归视频生成框架,支持多模态交互与低延迟流式外推。通过最小改动标准大语言模型(LLM),该框架接收包括音频、姿态和文本在内的多模态条件编码,并输出空间与语义一致的表征以指导扩散头的去噪过程。为此,我们从多个来源构建了约20,000小时的大规模对话数据集,涵盖丰富对话场景用于训练。此外,引入深度压缩自编码器,压缩比高达64×,有效缓解自回归模型的长时序推理负担。大量实验表明,在双人对话、多语言人类合成及交互式世界建模任务中,本方法在低延迟、高效率和细粒度多模态控制方面均具优势。

原文摘要 · Abstract (English)

Recently, interactive digital human video generation has attracted widespread attention and achieved remarkable progress. However, building such a practical system that can interact with diverse input signals in real time remains challenging to existing methods, which often struggle with heavy computational cost and limited controllability. In this work, we introduce an autoregressive video generation framework that enables interactive multimodal control and low-latency extrapolation in a streaming manner. With minimal modifications to a standard large language model (LLM), our framework accepts multimodal condition encodings including audio, pose, and text, and outputs spatially and semantically coherent representations to guide the denoising process of a diffusion head. To support this, we construct a large-scale dialogue dataset of approximately 20,000 hours from multiple sources, providing rich conversational scenarios for training. We further introduce a deep compression autoencoder with up to 64$\times$ reduction ratio, which effectively alleviates the long-horizon inference burden of the autoregressive model. Extensive experiments on duplex conversation, multilingual human synthesis, and interactive world model highlight the advantages of our approach in low latency, high efficiency, and fine-grained multimodal controllability.

数字人多模态自回归实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。