用一张照片生成能持续对话的数字人,支持实时音视频交互。
X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- 双模块统一架构,分别处理理解与生成,实现多模态同步输出。
- 单张照片可支撑数小时稳定对话,支持流式输入实时响应。
- 适合构建虚拟助手、数字偶像等需长期互动的AI角色。
我们提出X-Streamer,一个端到端的多模态人类世界建模框架,可构建能跨文本、语音和视频进行无限交互的数字人类代理。从单一肖像图出发,X-Streamer支持由流式多模态输入驱动的实时开放式视频通话。其核心为思考者-行动者双变压器架构,统一多模态感知与生成,将静态肖像转化为持续且智能的音视频互动。思考者模块感知并推理流式用户输入,其隐状态经行动者转换为实时同步的多模态流。具体而言,思考者利用预训练的大语言-语音模型,行动者采用分块自回归扩散模型,跨注意力机制结合思考者隐状态,生成时间对齐的多模态响应,包含交错的离散文本与音频标记及连续视频潜变量。为保障长时稳定性,设计了跨块与块内注意力,配合时间对齐的多模态位置编码,实现细粒度跨模态对齐与上下文保留,并通过分块扩散强制与全局身份引用进一步强化。X-Streamer可在两块A100 GPU上实时运行,从任意肖像图出发维持数小时一致的视频聊天体验,为交互式数字人类的统一世界建模开辟新路径。
原文摘要 · Abstract (English)
We introduce X-Streamer, an end-to-end multimodal human world modeling framework for building digital human agents capable of infinite interactions across text, speech, and video within a single unified architecture. Starting from a single portrait, X-Streamer enables real-time, open-ended video calls driven by streaming multimodal inputs. At its core is a Thinker-Actor dual-transformer architecture that unifies multimodal understanding and generation, turning a static portrait into persistent and intelligent audiovisual interactions. The Thinker module perceives and reasons over streaming user inputs, while its hidden states are translated by the Actor into synchronized multimodal streams in real time. Concretely, the Thinker leverages a pretrained large language-speech model, while the Actor employs a chunk-wise autoregressive diffusion model that cross-attends to the Thinker's hidden states to produce time-aligned multimodal responses with interleaved discrete text and audio tokens and continuous video latents. To ensure long-horizon stability, we design inter- and intra-chunk attentions with time-aligned multimodal positional embeddings for fine-grained cross-modality alignment and context retention, further reinforced by chunk-wise diffusion forcing and global identity referencing. X-Streamer runs in real time on two A100 GPUs, sustaining hours-long consistent video chat experiences from arbitrary portraits and paving the way toward unified world modeling of interactive digital humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。