arXiv:2608.21360cs.CV2026-08

构建首个面向多模态大模型的实时交互评估基准,测试其作为视频助手的实战能力。

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

论文配图:OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
图 1 · 摘自论文原文
  • 通过逆向工程网络视频生成多轮交互数据集,模拟真实人机协作场景。
  • 顶级模型Gemini-3-Pro得分66.4/100,开源模型Qwen3-Omni-Instruct得51.2,普遍表现不足。
  • 模型在手势识别、上下文记忆和响应时机上存在明显缺陷,亟待改进。

近期多模态大语言模型(Omni-LLMs)展现出作为实时视频助手的巨大潜力,能持续感知环境并引导用户达成目标。与传统被动视频理解不同,交互式助手需主动融合视觉状态、用户目标与先验知识提供有效帮助。评估极具挑战性,因模型不可预测的响应会动态改变用户后续行为,静态离线数据集难以应对。为此,我们提出OmniAssistBench。为解决同一目标可通过多种路径达成的问题,我们基于源视频提供预定义先验,要求模型引导用户走完全相同的路线。由于真实交互视频稀缺,我们通过逆向工程现有网络视频构建数据集,推断逻辑用户目标并将视频分割为多轮片段以模拟连续交互。该严格流程耗时超1000名专家小时。结果表明,专有模型Gemini-3-Pro得分为66.4/100,开源模型Qwen3-Omni-Instruct得51.2。尽管当前模型普遍理解用户输入,但常给出错误或不完整的回答。具体表现为:难以处理视觉提示(如手势)、多轮对话中无法保持历史上下文、未能延迟响应至目标事件发生。结果表明,模型距离成为可靠助手仍有巨大提升空间。

原文摘要 · Abstract (English)

Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.

多模态交互评估视频助手基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。