打造能自适应交互、自我进化的人格化数字人
Towards Interactive Intelligence for Digital Humans
- 五模块一体化框架实现认知与实时表现融合
- 在新基准上性能超越现有方法,多维度表现优异
- 适合虚拟助手、游戏角色等需要智能互动的场景
我们提出交互智能(Interactive Intelligence)这一数字人的新范式,具备人格化表达、自适应交互和自我演进能力。为此,我们构建了端到端的Mio(Multimodal Interactive Omni-Avatar)框架,包含五个专用模块:Thinker、Talker、Face Animator、Body Animator和Renderer。该统一架构将认知推理与实时多模态具身表现相结合,实现流畅一致的交互体验。此外,我们建立了一个新基准,用于严格评估交互智能能力。大量实验表明,本框架在所有评估维度上均优于当前最先进方法。这些贡献推动数字人从表面模仿走向真正智能交互。
原文摘要 · Abstract (English)
We introduce Interactive Intelligence, a novel paradigm of digital human that is capable of personality-aligned expression, adaptive interaction, and self-evolution. To realize this, we present Mio (Multimodal Interactive Omni-Avatar), an end-to-end framework composed of five specialized modules: Thinker, Talker, Face Animator, Body Animator, and Renderer. This unified architecture integrates cognitive reasoning with real-time multimodal embodiment to enable fluid, consistent interaction. Furthermore, we establish a new benchmark to rigorously evaluate the capabilities of interactive intelligence. Extensive experiments demonstrate that our framework achieves superior performance compared to state-of-the-art methods across all evaluated dimensions. Together, these contributions move digital humans beyond superficial imitation toward intelligent interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。