首个在线视频问答中断测试基准,评估模型实时响应打断能力。
OVIBench: Benchmarking Online Video Question Answering under Interruption

- 构建在线视频问答中断任务,模拟用户中途打断场景
- 三类中断类型+开闭卷评测,支持大规模可复现测试
- 训练数据与基准同步发布,适合提升模型交互鲁棒性
近期视觉语言模型在视频理解方面取得显著进展,但现有视频问答研究与基准大多遵循离线单轮范式,忽视了用户在模型生成过程中可能中断的真实交互场景。为此,我们提出在线视频问答中断任务,并引入首个标准化基准OVIBench,用于评估视觉语言模型在此设置下的表现。OVIBench将中断分为三类:取消、误触发、修正,并支持开放式与多选题评估。为实现大规模可复现测试,我们设计了离线模拟协议,在统一时间框架下重现生成过程中的中断行为,并提供多维评估指标以衡量中断理解与响应生成能力。实验表明,OVIBench能有效区分模型的中断处理能力,尤其在修正请求方面表现明显。最后,我们构建了用于中断感知微调的训练集OVI-Train。在该数据集上微调的模型在OVIBench上取得显著提升,验证了基准与数据设计的有效性。OVIBench、OVI-Train及评估代码将公开发布。
原文摘要 · Abstract (English)
Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting. OVIBench categorizes interruptions into three types: Cancellation, False Trigger, Correction and supports both open-ended and multiple-choice evaluations. To enable large-scale and reproducible testing, we develop an offline simulation protocol that reproduces interruption during generation under a unified temporal setup, together with a multi-dimensional metric suite for assessing interruption understanding and response generation. Experiments demonstrate that OVIBench effectively distinguishes models' interruption-handling abilities, especially in following correction requests. Finally, we construct a train set OVI-Train for interruption-aware fine-tuning. Models fine-tuned on this dataset achieve significant gains on OVIBench, validating the effectiveness of our benchmark and data design. OVIBench, OVI-Train, and the evaluation code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。