MiniCPM-o 4.5实现边看边听边说的实时全双工多模态交互。
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

- 采用统一时间轴的Omni-Flow框架,实现多模态输入输出同步处理。
- 在90亿参数下性能接近Gemini 2.5 Flash,且推理效率更高。
- 可在12GB内存设备上运行,适合边缘部署,支持主动提醒等行为。
多模态大模型虽已从静态处理迈向实时流交互,但仍远未达到人类水平。核心瓶颈在于交互范式:感知与响应仍分时交替,无法实时调整;多数模型被动响应,缺乏主动行为。我们提出MiniCPM-o 4.5,通过实时全双工多模态交互,实现同时看、听、说,并能基于对实时场景的理解主动发出提醒或评论。其核心技术是Omni-Flow,一种将多模态输入输出沿共享时间轴对齐的统一流式框架,将传统轮次交互转化为时间对齐的全双工过程。该模型共90亿参数,在视觉-语言能力上接近Gemini 2.5 Flash,开源表现领先同规模模型;在多模态理解上超越Qwen3-Omni-30B-A3B,语音生成更优,计算效率显著提升。得益于高效架构与推理优化,可在小于12GB RAM的边缘设备上实现实时交互。
原文摘要 · Abstract (English)
Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。