评测视觉语言模型对长视频质量的感知理解能力,发现越长越难准确判断。
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

- 设计三层次评估框架,覆盖局部、跨事件到全局质量理解
- 包含1200段多类型长视频和1500个问题,测试模型长期推理能力
- 引入稀疏扰动检测任务,检验模型对细微失真细节的识别水平
大型视觉语言模型(LVLMs)在长期视频质量理解方面的评估仍面临挑战。现有视频质量基准主要集中于短片段和孤立失真,忽视了长时内容中的时间连续性、累积退化及推理复杂性。为此,我们提出LongVQUBench,一个全面的长期视频质量理解基准。该基准包含超过1200段多样化的视频,涵盖电影、纪录片、监控录像、第一人称记录和动画内容,并配有1500个选择题与开放问答用于验证与测试。为评估不同时间尺度下的感知推理能力,我们引入三个逐步复杂的评估层级:(i) 局部事件质量理解(LQU),分析局部失真;(ii) 跨事件质量推理(CQR),整合多个退化事件;(iii) 全局质量理解(GQU),对长时间段进行整体感知评估。此外,在所有层级中嵌入针状失真问答(NDQA)范式,通过稀疏插入空间或时间伪影,探测模型对细粒度失真检测与推理的能力。对14个先进LVLM的大量实验显示,随着视频长度增加和推理深度加深,性能显著下降,凸显其在长程时间整合与感知归因上的局限性。我们期望LongVQUBench成为系统化、分层化、可解释的LVLM长期视频质量理解评估的基石。
原文摘要 · Abstract (English)
The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。