统一生成数字人多模态内容,支持文本音频动作视觉协同输出
Archon: A Unified Multimodal Model for Holistic Digital Human Generation

- 用专用分词器统一七种模态,自回归建模联合分布
- 视频重参数化降4倍令牌数,保持精细动态细节
- 跨模态任务分步思维链提升生成可控性与质量
数字人是沉浸式交互的核心,但统一生成文本、音频、动作和视觉内容的模型仍具挑战。本文提出 Archon,一个全预训练、以人类为中心的统一多模态模型,用于整体数字人生成。Archon 使用七种模态专用分词器,并基于同步多模态数据和72个多样化任务进行预训练,建模完整的联合分布。为解决高保真对话视频中的令牌爆炸问题,提出一种内存高效的语义视频重参数化方法,在保留细粒度动态的前提下实现4倍令牌压缩,并结合语义驱动的视频扩散解码器。进一步提出“在模态中思考”机制,将模糊的跨模态任务分解为交替模态的逐步推理链,逐步提升生成质量和可控性。大量实验表明,Archon 在多种数字人生成任务中表现优异或相当,验证了该统一框架的有效性。
原文摘要 · Abstract (English)
Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human-centric unified multimodal model for holistic avatar generation. Archon unifies seven modalities with modality-specific tokenizers, and a native autoregressive unified multimodal model pretrained on synchronized modalities and 72 diverse tasks to model holistic joint distributions. To address the token explosion challenge in high-fidelity talking videos, we introduce a memory-efficient semantic video reparameterization, achieving 4x token reduction while preserving fine-grained dynamics, coupled with a semantic-driven video diffusion decoder. We further propose a "Thinking in Modality" that decomposes ambiguous cross-modal tasks into stepwise thinking in an alternative chain of modality, progressively enhancing fidelity and controllability. Extensive experiments demonstrate that Archon achieves superior or comparable performance across diverse digital human generation tasks, validating the effectiveness of our unified framework. Project page: https://zju3dv.github.io/archon/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。