首个评估多模态大模型流式主动智能的基准,揭示其在连续交互中的短板。
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

- 构建流式视频场景下的主动智能评估基准,涵盖多轮交互与混合请求
- 发现主流模型存在主动触发不稳定、反应与主动行为协调弱的问题
- 提出无需训练的代理框架,通过控制策略和时间门机制提升稳定性
近期多模态大语言模型(MLLMs)在被动问答任务中表现优异,但现实中的流式助手需对持续视觉输入进行主动推理。现有基准主要关注孤立的单轮反应或主动交互,忽视了用户可随时增删或修改主动请求,并与反应式查询交错的动态多轮场景。为此,我们提出 IPIBench,首个面向流式视频场景下 MLLMs 交互式主动智能的评估基准。该基准覆盖主动监控、主动任务管理及交错的反应-主动请求。对代表性 MLLMs 的评估揭示两大局限:主动触发不稳定,反应与主动行为间协调性差。我们进一步提出 IPI-Agent,一种无需训练的代理框架,包含交互控制策略与时间门机制,以稳定主动触发并协调多轮交互。实验表明,IPI-Agent 在所有基准设置下均一致提升现有 MLLMs 性能。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over continuous visual inputs. Existing benchmarks mainly study reactive or proactive interactions in isolated single-turn settings, overlooking dynamic multi-turn scenarios where users may add, modify, or cancel proactive requests alongside interleaved reactive queries. To address this gap, we introduce IPIBench, the first benchmark for evaluating Interactive Proactive Intelligence of MLLMs under streaming video settings. IPIBench covers proactive monitoring, proactive task management, and interleaved reactive-proactive requests. Evaluations on representative MLLMs reveal two major limitations: unstable proactive triggering and weak coordination between reactive and proactive behaviors. We further propose IPI-Agent, a training-free agentic framework with an interaction-control policy and a temporal-gating mechanism for stabilizing proactive triggering and coordinating multi-turn interactions. Experiments show that IPI-Agent consistently improves existing MLLMs across all benchmark settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。