用规则引导的伪标签提升零样本视频摘要质量
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
- 将少量人工标注转为高置信度伪标签,构建适应数据集的评分规则
- 边界帧独立评分,中间帧结合邻近片段摘要评估连贯性与冗余
- 无需训练即可实现通用与查询聚焦摘要,结果稳定且可解释
我们提出一种基于评分规则引导、伪标签生成和提示驱动的零样本视频摘要框架,将大语言模型与结构化语义推理相结合。将少量人工标注转化为高置信度伪标签,并组织为适应数据集的评分规则,明确涵盖主题相关性、动作细节和叙事进展等维度。推理时,边界场景(如开头结尾)基于自身描述独立评分,中间场景则结合相邻片段的简要摘要,评估叙事连贯性与冗余性。该设计使语言模型在无参数微调情况下平衡局部显著性与全局一致性。在三个基准上表现稳定且优异:SumMe 上 F1 达 57.58,TVSum 上达 63.05,QFVS 上达 53.79,分别优于零样本基线 +0.85、+0.84 和 +0.37。结果表明,规则引导的伪标签结合上下文提示,有效稳定了基于 LLM 的评分,建立了一种通用、可解释、无需训练的视频摘要范式。
原文摘要 · Abstract (English)
We propose a rubric-guided, pseudo-labeled, and prompt-driven zero-shot video summarization framework that bridges large language models with structured semantic reasoning. A small subset of human annotations is converted into high-confidence pseudo labels and organized into dataset-adaptive rubrics defining clear evaluation dimensions such as thematic relevance, action detail, and narrative progression. During inference, boundary scenes, including the opening and closing segments, are scored independently based on their own descriptions, while intermediate scenes incorporate concise summaries of adjacent segments to assess narrative continuity and redundancy. This design enables the language model to balance local salience with global coherence without any parameter tuning. Across three benchmarks, the proposed method achieves stable and competitive results, with F1 scores of 57.58 on SumMe, 63.05 on TVSum, and 53.79 on QFVS, surpassing zero-shot baselines by +0.85, +0.84, and +0.37, respectively. These outcomes demonstrate that rubric-guided pseudo labeling combined with contextual prompting effectively stabilizes LLM-based scoring and establishes a general, interpretable, and training-free paradigm for both generic and query-focused video summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。