AURA让视频大模型实时理解直播画面并即时回答问题。
AURA: Always-On Understanding and Real-Time Assistance via Video Streams

- 构建端到端流式框架,持续处理视频流并支持实时问答
- 在2帧/秒下实现语音识别与语音合成,双80G显卡运行
- 适合需要连续交互的智能助手、监控系统等场景
视频大语言模型在多项视频理解任务中表现优异,但现有系统多为离线处理,难以应对需持续观察和及时响应的实时视频流。近期流式视频模型有所进展,但多数依赖分离的触发-响应流程,或仅限于生成描述性文字,限制了开放问题问答与长时交互能力。本文提出AURA(Always-On Understanding and Real-Time Assistance),一种端到端流式视觉交互框架,使统一的视频大模型能持续处理视频流,并支持实时问答与主动响应。AURA整合了上下文管理、数据构造、训练目标与部署优化,确保长时间交互的稳定性。在流式基准上达到领先性能,支持实时演示系统,结合语音识别与语音合成,在两块80G显卡上以2帧/秒运行。我们开源了AURA模型与实时推理框架,以推动后续研究。
原文摘要 · Abstract (English)
Video Large Language Models (VideoLLMs) have achieved strong performance on many video understanding tasks, but most existing systems remain offline and are not well-suited for live video streams that require continuous observation and timely response. Recent streaming VideoLLMs have made progress, yet current approaches often rely on decoupled trigger-response pipelines or are limited to captioning-style narration, reducing their effectiveness for open-ended question answering and long-horizon interaction. We propose AURA (Always-On Understanding and Real-Time Assistance), an end-to-end streaming visual interaction framework that enables a unified VideoLLM to continuously process video streams and support both real-time question answering and proactive responses. AURA integrates context management, data construction, training objectives, and deployment optimization for stable long-horizon streaming interaction. It achieves state-of-the-art performance on streaming benchmarks and supports a real-time demo system with ASR and TTS running at 2 FPS on two 80G accelerators. We release the AURA model together with a real-time inference framework to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。