用视频属性评估视频问答对复杂度,提升大模型评测准确性。
An Attribute-Based Measure of Video Complexity

- 基于视频属性空间的非参数化复杂度测量方法。
- 在小样本下仍优于现有评测方式,准确率显著更高。
- 可解释性强,适合分析评测基准设计缺陷。
本文提出一种新的视频-语言大模型(video-LLM)复杂度评估框架——视频属性基复杂度(VideoABC),将复杂度定义为特定视频-问题对导致video-LLM失败的概率。VideoABC采用参考视频数据集与预定义视频属性词典(如场景复杂度、事件速度)构建非参数化度量。训练阶段将参考视频投影至属性空间并量化,计算每个量化单元的期望复杂度。给定新视频,通过其属性投影定位对应单元,估算复杂度。为应对小样本参考集,融合k均值量化器与通用格量化器:前者保证分布内样本精度,后者确保分布外泛化能力。通过受心理物理学启发的合成视频生成方法填充格量化器单元,实现期望复杂度计算。实验表明,即使使用低维属性表示,VideoABC仍显著优于‘video-LLM作为评判者’等方法,且其可解释性揭示了评测基准属性组成对复杂度的影响。
原文摘要 · Abstract (English)
A new framework for the estimation of the complexity posed by video-question pairs to video-LLMs, Video Attribute-Based Complexity (VideoABC), is proposed. Video complexity is defined as the probability of failure of a video-LLM for a given video-question pair. VideoABC is a non-parametric complexity measure, using a reference video dataset and a pre-defined vocabulary of video attributes informative of complexity, \eg the scene complexity or the speed of the video event informative of the question. In a training phase, reference videos are projected into the space of these attributes, which is then quantized. The expected ABC of each quantization cell is then computed. Given a new video and its projection into the attribute space, complexity is estimated by the expected ABC of the associated quantization cell. To enable the use of VideoABC with small reference video datasets, two quantizers are combined: a k-means quantizer that enables accurate complexity estimates for samples in the distribution of the reference dataset and a universal lattice quantizer that guarantees generalization to out-of-distribution samples. A synthetic video generation procedure, inspired by target-distractor manipulations of psychophysics studies, is proposed to populate the cells of the lattice quantizer during training, enabling the computation of their expected ABCs. Experimental results show that VideoABCis effective even with very low-dimensional attribute representations, substantially outperforming approaches like `video-LLM as judge' with much less complexity. Finally, the explainable nature of the VideoABC score, in terms of well-defined attributes, is shown to provide insights on how the attribute composition of benchmarks affects their complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。