让AI像人一样实时感知并主动交互,无需等待指令。
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

- 8B模型持续观看视频,自主决定何时回应或调用后台模型。
- 在6个真实场景中,人类评分显著优于百度、Gemini的对话助手。
- 开源完整系统,支持语音、记忆、界面等模块灵活接入。
现实世界中的许多事件不会等待用户提问:监控画面突发火灾、视频通话中表情一闪而过、直播中商品快速闪过。然而当前大模型仍以问答轮次为主,即使视频通话应用也仅在被触发时响应。本文提出新范式:让模型如人般持续感知、自主判断是否发声,并实现实时交互。我们发布8B规模的视觉优先多模态交互模型JoyAI-VL-Interaction,每秒自主决策沉默、响应或调用背景模型,擅长视觉触发响应与时间感知。配套提供可迁移训练方案,使其具备未显式训练的能力,如引导用户切换应用界面或根据幻灯片即兴讲授。同时开源完整部署系统,可实时流式处理任意视频输入,支持语音识别、文本转语音、记忆模块、可视化界面及外部API/智能体连接。在六个真实场景中,人类评估者对JoyAI-VL-Interaction的偏好远超百度和Gemini的内置助手。据我们所知,这是首个同时开放模型、训练方法、数据与完整可部署系统的视觉驱动交互模型。
原文摘要 · Abstract (English)
Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes by in a livestream. Yet today's large models remain mostly turn-based by design: they answer only when addressed, and even video-call apps that appear interactive still operate as question-answer systems, reacting only when polled or prompted. We argue for a different paradigm: a model that is present in the world like a person. It continuously watches what is happening now, decides on its own whether to speak or stay silent, interacts in real time, and delegates to a background model when the problem is hard. To advance interaction models and their adoption across domains, we make two fully open-sourced contributions. First, we release JoyAI-VL-Interaction, an 8B-scale, vision-first VL-interaction model. The model makes the response decision internally, choosing each second to stay silent, respond, or delegate to a background model, and it excels at vision-triggered responsiveness and time awareness. We pair it with a transferable training recipe, from which capabilities we never trained for emerge, such as guiding a shopper through changing app screens or improvising a lecture from a slide deck. Second, we release a complete, deployable system built around that model. The system streams any ongoing video into the model, making it genuinely present in the world. All other components are pluggable, including ASR/TTS modules, memory, visualization UI, and a background brain that can connect to any API or agent. Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin. To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。