arXiv:2506.05395cs.CVcs.IR2025-06被引 3

用三模态信息精准提取视频关键帧,提升摘要与检索效果。

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations

  • 融合色彩、结构、语义三类特征,生成紧凑多模态嵌入
  • 在TVSum20和SumMe上超越现有方法,性能领先
  • 适合需要高质量视频摘要的多媒体应用

高效的关键帧提取对视频摘要与检索至关重要,但完整捕捉视频的语义与视觉丰富性仍具挑战。我们提出TriPSS,一种融合感知(CIELAB颜色空间)、结构(ResNet-50)与语义(LLaMA-3.2-11B-Vision-Instruct生成的帧级描述)三模态特征的框架。通过主成分分析融合多模态表示,结合HDBSCAN实现自适应视频分割。经质量评估与重复帧过滤的优化阶段,确保关键帧集简洁且语义多样。在TVSum20与SumMe基准测试中,TriPSS表现卓越,显著优于单模态及已有多模态方法。结果表明其能有效捕获互补的视觉与语义线索,是视频摘要、检索与大规模多媒体理解的有效方案。

原文摘要 · Abstract (English)

Efficient keyframe extraction is critical for video summarization and retrieval, yet capturing the full semantic and visual richness of video content remains challenging. We introduce TriPSS, a tri-modal framework that integrates perceptual features from the CIELAB color space, structural embeddings from ResNet-50, and semantic context from frame-level captions generated by LLaMA-3.2-11B-Vision-Instruct. These modalities are fused using principal component analysis to form compact multi-modal embeddings, enabling adaptive video segmentation via HDBSCAN clustering. A refinement stage incorporating quality assessment and duplicate filtering ensures the final keyframe set is both concise and semantically diverse. Evaluations on the TVSum20 and SumMe benchmarks show that TriPSS achieves state-of-the-art performance, significantly outperforming both unimodal and prior multimodal approaches. These results highlight TriPSS' ability to capture complementary visual and semantic cues, establishing it as an effective solution for video summarization, retrieval, and large-scale multimedia understanding.

视频摘要多模态关键帧提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。