用轻量模型自动识别体育赛事精彩片段,准确率超80%。
Automated Detection of Sport Highlights from Audio and Video Sources
- 基于音频梅尔频谱和灰度视频帧训练小模型,实现快速部署。
- 音频检测准确率达89%,视频达83%,融合模态效果更优。
- 适合需要低成本自动剪辑的体育内容平台或媒体机构。
本研究提出一种基于深度学习的轻量化方法,用于从音视频源中自动检测体育赛事精彩片段(HLs)。HL检测是体育视频分析的关键任务,传统方式依赖大量人工。该方法在较小规模的音频梅尔频谱与灰度视频帧数据集上训练,音频与视频检测准确率分别达到89%和83%。采用简单架构与小数据集,证明了其在快速、低成本部署上的可行性。此外,融合音视频双模态的集成模型显著提升了抗误检与漏检的能力。该方法可扩展至多种体育内容的自动化处理,减少人工干预。未来工作将优化模型结构,并拓展至更广泛的媒体场景检测任务。
原文摘要 · Abstract (English)
This study presents a novel Deep Learning-based and lightweight approach for the automated detection of sports highlights (HLs) from audio and video sources. HL detection is a key task in sports video analysis, traditionally requiring significant human effort. Our solution leverages Deep Learning (DL) models trained on relatively small datasets of audio Mel-spectrograms and grayscale video frames, achieving promising accuracy rates of 89% and 83% for audio and video detection, respectively. The use of small datasets, combined with simple architectures, demonstrates the practicality of our method for fast and cost-effective deployment. Furthermore, an ensemble model combining both modalities shows improved robustness against false positives and false negatives. The proposed methodology offers a scalable solution for automated HL detection across various types of sports video content, reducing the need for manual intervention. Future work will focus on enhancing model architectures and extending this approach to broader scene-detection tasks in media analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。