arXiv:2602.23739cs.CV2026-02中稿 · CVPR被引 1

首个支持实时多模态交互的统一系统,让虚拟助手能同步说话、动作和视觉反馈。

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

  • 通过分段对齐策略实现多模态同步生成
  • 采用重演驱动学习保持推理能力,性能优于现有方法
  • 适合开发智能对话机器人、虚拟形象等沉浸式应用

构建具备自然动态交互能力的智能体,需实现全栈实时多模态交互。然而现有系统或仅支持单模态生成,或因推理能力下降与跨模态对齐差,导致交互不连贯。本文提出U-Mind,首个支持实时生成的统一多模态对话系统,可在单一交互循环中联合建模语言、语音、动作与视频合成。核心为统一对齐与推理框架:通过分段对齐策略提升跨模态同步性,借助重演驱动学习保持推理能力。推理时采用文本优先解码流程,先进行内部思维链规划,再实现多模态时序同步生成。同时设计基于姿态与语音的实时视频渲染框架,实现具表现力的同步视觉反馈。大量实验表明,U-Mind在问答、指令遵循与动作生成等任务上达到当前最优性能,推动智能沉浸式对话代理的发展。

原文摘要 · Abstract (English)

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coherent and perceptually grounded interactions. In this work, we introduce U-Mind, the first unified system for high-intelligence multimodal dialogue that supports real-time generation and jointly models language, speech, motion, and video synthesis within a single interactive loop. At its core, U-Mind implements a Unified Alignment and Reasoning Framework that addresses two key challenges: enhancing cross-modal synchronization via a segment-wise alignment strategy, and preserving reasoning abilities through Rehearsal-Driven Learning. During inference, U-Mind adopts a text-first decoding pipeline that performs internal chain-of-thought planning followed by temporally synchronized generation across modalities. To close the loop, we implement a real-time video rendering framework conditioned on pose and speech, enabling expressive and synchronized visual feedback. Extensive experiments demonstrate that U-Mind achieves state-of-the-art performance on a range of multimodal interaction tasks, including question answering, instruction following, and motion generation, paving the way toward intelligent, immersive conversational agents.

多模态交互实时生成虚拟形象统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。