无需人工标注,自动筛选高质量视频摘要。
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
- 基于信息瓶颈原理,用两个指标评估摘要质量。
- 在多个数据集上提升任务准确率最高达61.23%。
- 适合需要快速决策的视频审查场景。
许多既需准确性又需效率的决策任务仍依赖人工干预,例如交通员审看长达一小时的行车记录仪视频或研究人员筛选会议视频。现有视觉语言模型生成的摘要往往冗长重复,影响任务表现。当前视频摘要评价依赖昂贵的人工标注,且忽略摘要在下游任务中的实用性。本文提出无需人工标注的视频到文本信息瓶颈评估方法(VIBE),通过两个指标——对齐度(摘要与视觉内容的一致性)和效用(对任务的帮助程度)——评分并筛选最优摘要。在LearningPaper24、SUTD-TrafficQA和LongVideoBench上的人类实验表明,VIBE选出的摘要相比原始模型输出或原始视频,可使任务准确率最高提升61.23%,响应时间减少75.77%。
原文摘要 · Abstract (English)
Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance-boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。