打造能实时双向交互的全模态机器人模型,响应快到80毫秒。
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- 用统一架构实现视觉、语音、文本等多模态实时处理
- 全双工延迟仅80毫秒,对话自然度显著提升
- 适合开发真实场景下的智能人机交互系统
人类天然以全双工方式处理现实世界的多模态信息。在人工智能领域,复现这一能力对推动模型研发与部署至关重要,尤其在具身智能场景中。当前多模态模型面临两大挑战:(1)有效处理超过三种模态(如视觉、音频、文本);(2)对快速变化的人类指令做出全双工响应。为推动支持全模态处理与全双工能力的模型研究,我们提出RoboEgo(又称FLM-Ego),一个原生支持全双工的统一模型系统。该系统采用具备全双工能力的骨干架构与算法,理论双工延迟低至80毫秒。在真实环境下流式视觉语境对话中,RoboEgo展现出更优的响应速度与语音自然度,同时内容质量与最先进半双工全模态模型相当——这一成果此前被认为原生全双工系统无法实现。
原文摘要 · Abstract (English)
Humans naturally process real-world multimodal information in a full-duplex manner. In artificial intelligence, replicating this capability is essential for advancing model development and deployment, particularly in embodied contexts. The development of multimodal models faces two primary challenges: (1) effectively handling more than three modalities-such as vision, audio, and text; and (2) delivering full-duplex responses to rapidly evolving human instructions. To facilitate research on models that support both omnimodal processing and full duplexity, we present RoboEgo (alias: FLM-Ego), a unified model system designed to address both challenges. RoboEgo incorporates a backbone architecture and algorithms that natively support full duplexity, achieving a theoretical duplex latency of 80 ms. In streaming visually grounded conversations under real-world conditions, RoboEgo exhibits superior responsiveness and speech naturalness, while maintaining comparable content qualities to state-of-the-art semi-duplex omnimodal models-a feat previously considered unattainable by native full-duplex systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。