arXiv:2605.09507cs.CV2026-05中稿 · presentation at th…

提出新框架,让视频摘要更贴近人类偏好且推理更快。

Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

论文配图:Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
图 1 · 摘自论文原文
  • 用概率分数建模不确定性,更好处理多人标注差异。
  • 在SumMe和TVSum上相关性指标优于传统方法。
  • 适合需要高效、稳定摘要的场景,如视频推荐系统。

视频摘要旨在通过选取时间上重要的片段,生成长视频的紧凑表示,反映人类偏好。该任务因标注主观性强及评估依赖离散解码(如时间分段与背包选择)而极具挑战。现有方法或仅学习确定性重要性分数,忽略主观性;或采用复杂生成模型,增加训练与推理开销。本文提出VASTSum框架,通过变分形式预测帧级概率重要性分数,显式建模多标注者监督带来的不确定性。针对二值标注下的主观性,采用鼓励对齐合理人类标注模式的监督策略,而非强制单一共识目标。此外引入解码对齐正则化,提升背包式摘要选择的稳定性,降低对预测分数微小扰动的敏感性。在SumMe和TVSum基准上使用标准秩相关指标评估,实验显示在多个数据划分下,肯德尔与斯皮尔曼相关性持续且领先,证明在标注不一致情况下仍具鲁棒性,同时保持单次前向传播的高效推理。结果表明,显式建模不确定性并使学习目标与解码阶段对齐,为视频摘要提供了优于确定性与扩散模型的合理范式。

原文摘要 · Abstract (English)

Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect human preferences. This task is inherently difficult due to strong annotation subjectivity and the reliance on discrete decoding procedures, such as temporal segmentation and knapsack-based selection, during evaluation. Most existing approaches either learn deterministic importance scores that overlook these characteristics or adopt complex generative models that increase training and inference cost. In this paper, we propose VASTSum, an uncertainty-aware and decoder-aligned learning framework for video summarization that addresses both challenges within a single-pass model. The proposed method predicts probabilistic frame-level importance scores using a variational formulation, enabling explicit modeling of uncertainty arising from multi-annotator supervision. To account for subjectivity, particularly under binary annotations, we employ a supervision strategy that encourages alignment with plausible human annotation modes rather than enforcing a single consensus target. Furthermore, we introduce a decoder-aligned regularization that promotes stability of knapsack-based summary selection, reducing sensitivity to small perturbations in predicted scores. We evaluate the proposed framework on the SumMe and TVSum benchmarks using standard rank-based metrics. Experimental results show consistent and competitive Kendall and Spearman correlations across multiple data splits, demonstrating improved robustness under annotation disagreement while maintaining efficient single-forward inference. These results indicate that explicitly modeling uncertainty and aligning learning objectives with the decoding stage provide a principled alternative to both deterministic and diffusion-based video summarization methods.

视频摘要不确定性建模解码对齐高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。