让AI像人一样主动看视频并及时回答问题
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- 设计新框架与数据集,让模型能主动理解视频流
- 在多个测试中表现优于现有模型,响应更及时准确
- 适合做实时视频助手或智能监控系统的开发者
设想一种能在人类场景中运行的AI,不仅能观察,还能主动理解、预判并适时回应动态事件。为此,我们提出一个创新任务:给定第一人称视角视频流,助手需在恰当时刻主动回答不断变化的问题,同时保持感知与推理同步。该任务包含三大特性:(1) 主动连贯性,(2) 即时响应性,(3) 同步高效性。为评估这些特性,我们首次引入ESTP-Bench(第一人称主动基准)和ESTP-F1指标,构建全新评估体系。同时提出完整技术方案,包括:(1) 数据引擎,(2) 多阶段训练策略,(3) 主动动态压缩技术。所提模型有效解决上述关键挑战,在多种在线与离线基准上超越多个基线。项目页面:https://zhangyl4.github.io/publications/eyes-wide-open/
原文摘要 · Abstract (English)
Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given ego-streaming video input, an assistant proactively answers diverse, evolving questions at the opportune moment, while maintaining synchronized perception and reasoning. This task embodies three key properties: (1) Proactive Coherence, (2) Just-in-Time Responsiveness, and (3) Synchronized Efficiency. To evaluate and address these properties, we first introduce ESTP-Bench (Ego Streaming Proactive Benchmark) alongside the ESTP-F1 metric-a novel framework designed for their rigorous assessment. Secondly, we propose a comprehensive technical pipeline to enable models to tackle this challenging task. This pipeline comprises: (1) a data engine, (2) a multi-stage training strategy, and (3) a proactive dynamic compression technique. Our proposed model effectively addresses these critical properties while outperforming multiple baselines across diverse online and offline benchmarks. Project Page:https://zhangyl4.github.io/publications/eyes-wide-open/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。