评测视觉语言模型在长视频理解中的可信度,防止盲目猜测误导评估结果。
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
- 构建多采样帧级的视频评测集,区分可答与不可答问题。
- 顶尖模型拒答准确率达70%以上,最差模型接近0%。
- 多数模型不被明确要求时拒绝回答意愿大幅下降,适合关注模型可靠性研究者。
近期视觉-语言模型(VLMs)在多模态理解任务中取得显著进展,但其在长视频理解上的评估仍不可靠。由于输入帧数量有限,关键帧可能缺失,导致模型无法作答。然而,那些因不确定而拒绝回答的模型会被标记为错误,而盲目猜测的模型可能偶然答对,从而获得虚高准确率,造成误导性评估结果,促使模型倾向于猜测而非诚实回应。为此,我们提出VirtueBench,一个专门评估模型在不确定性下可信度的基准。该基准为每段视频设置多个帧采样层级,并提供标注以区分可答与不可答情况。对25个开源及商用VLM的评估显示,不同模型家族的拒答行为差异显著,拒答准确率从最佳模型的70%以上到最差模型接近0%不等。此外,当提示未明确要求拒答时,多数模型的拒答意愿显著下降。这些发现凸显了在评估中引入可靠性和可信度导向的基准与排行榜的必要性。
原文摘要 · Abstract (English)
Recent Vision-Language Models (VLMs) have made remarkable progress in multimodal understanding tasks, yet their evaluation on long video understanding remains unreliable. Due to limited frame inputs, key frames necessary for answering the question may be missing from the model's input. However, models that truthfully refuse to answer under such uncertainty are marked as incorrect, while those that guess may coincidentally produce the correct answer and thus obtain deceptively higher accuracy, leading to misleading evaluation results and encouraging models to guess rather than respond honestly. To address this issue, we introduce VirtueBench, a benchmark explicitly designed to assess model trustworthiness under uncertainty. VirtueBench constructs multiple frame-sampling levels for each video and provides ground truths that distinguish between answerable and unanswerable cases. Evaluations on 25 open-source and commercial VLMs reveal distinct refusal behaviors across different model families, with refusal accuracy ranging from over 70% in the best models to nearly 0% in the worst. Moreover, most models exhibit a substantial drop in refusal when the prompt does not explicitly require them to do so. These findings highlight the need for developing trustworthy VLMs for multimodal understanding, guided by benchmarks and leaderboards that emphasize reliability and trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。