arXiv:2506.19168cs.CV2025-06

用视觉感知差异找视频关键帧,无需训练且实时高效。

PRISM: Perceptual Recognition for Identifying Standout Moments in Human-Centric Keyframe Extraction

  • 在CIELAB色彩空间中用感知色差识别显著帧
  • 在4个数据集上达高保真度与高压缩比
  • 适合实时监控和资源受限平台使用

在线视频在塑造政治话语和放大网络社会威胁(如虚假信息、宣传与极端化)中起核心作用。识别视频中最具影响力的‘突出时刻’对内容审核、摘要生成和取证分析至关重要。本文提出PRISM(Perceptual Recognition for Identifying Standout Moments),一种轻量级且符合人类视觉感知的关键词提取框架。PRISM基于CIELAB色彩空间,利用感知色差度量识别与人眼敏感度一致的帧。相比深度学习方法,PRISM无需训练、可解释性强且计算效率高,适用于实时与资源受限环境。我们在BBC、TVSum、SumMe和ClipShots四个基准数据集上评估,结果表明其在保持高压缩率的同时,仍具备优异的准确性和保真度。该表现证明了PRISM在结构化与非结构化视频内容中的有效性,具备作为可扩展工具用于在线平台有害或政治敏感媒体分析与管控的潜力。

原文摘要 · Abstract (English)

Online videos play a central role in shaping political discourse and amplifying cyber social threats such as misinformation, propaganda, and radicalization. Detecting the most impactful or "standout" moments in video content is crucial for content moderation, summarization, and forensic analysis. In this paper, we introduce PRISM (Perceptual Recognition for Identifying Standout Moments), a lightweight and perceptually-aligned framework for keyframe extraction. PRISM operates in the CIELAB color space and uses perceptual color difference metrics to identify frames that align with human visual sensitivity. Unlike deep learning-based approaches, PRISM is interpretable, training-free, and computationally efficient, making it well suited for real-time and resource-constrained environments. We evaluate PRISM on four benchmark datasets: BBC, TVSum, SumMe, and ClipShots, and demonstrate that it achieves strong accuracy and fidelity while maintaining high compression ratios. These results highlight PRISM's effectiveness in both structured and unstructured video content, and its potential as a scalable tool for analyzing and moderating harmful or politically sensitive media in online platforms.

视频摘要关键帧提取感知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。