图像分类器比视频音频模型更高效地分割新闻视频内容
Comparative Analysis of Image, Video, and Audio Classifiers for Automated News Video Segmentation
- 用图像模型替代复杂时序模型进行新闻视频分段
- 图像分类准确率达84.34%,优于复杂视频模型
- 适合媒体归档与智能搜索场景的轻量级部署
新闻视频需高效的內容组织与检索系统,但其非结构化特性给自动化处理带来挑战。本文对图像、视频和音频分类器在新闻视频自动分段中的表现进行了全面比较。基于自标注的41个新闻视频数据集(共1,832个片段),评估了ResNet、ViViT、AST及多模态架构等多种深度学习方法,用于识别广告、新闻故事、演播室场景、转场和可视化五类片段。实验表明,图像分类器性能最优(准确率84.34%),其中ResNet在远低于先进视频模型的计算开销下实现更高精度。二分类模型在转场(94.23%)和广告识别(92.74%)上表现尤为出色。研究为新闻视频分段提供了有效架构选择,并为媒体归档、个性化内容推送与智能视频搜索等应用提供实践指导。
原文摘要 · Abstract (English)
News videos require efficient content organisation and retrieval systems, but their unstructured nature poses significant challenges for automated processing. This paper presents a comprehensive comparative analysis of image, video, and audio classifiers for automated news video segmentation. This work presents the development and evaluation of multiple deep learning approaches, including ResNet, ViViT, AST, and multimodal architectures, to classify five distinct segment types: advertisements, stories, studio scenes, transitions, and visualisations. Using a custom-annotated dataset of 41 news videos comprising 1,832 scene clips, our experiments demonstrate that image-based classifiers achieve superior performance (84.34\% accuracy) compared to more complex temporal models. Notably, the ResNet architecture outperformed state-of-the-art video classifiers while requiring significantly fewer computational resources. Binary classification models achieved high accuracy for transitions (94.23\%) and advertisements (92.74\%). These findings advance the understanding of effective architectures for news video segmentation and provide practical insights for implementing automated content organisation systems in media applications. These include media archiving, personalised content delivery, and intelligent video search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。