通过音视频对齐实现多模态内容摘要,提升信息提取效率。
A Cascaded Architecture for Extractive Summarization of Multimedia Content via Audio-to-Text Alignment
- 分层架构结合语音转写与文本摘要模型,实现音视频内容自动提炼。
- 在ROUGE和F1指标上优于传统方法,即使存在转录误差仍表现稳定。
- 适合需要快速获取视频关键信息的研究者与内容平台用户。
本研究提出一种基于音频-文本对齐的多模态内容抽取式摘要级联架构。针对YouTube等多媒体资源中的关键信息提取难题,系统融合Microsoft Azure Speech进行语音转写,并结合Whisper、Pegasus、Facebook BART XSum等先进抽取式摘要模型。采用Pytube、Pydub与SpeechRecognition工具完成内容获取、音频提取与转录。通过命名实体识别与语义角色标注增强语言分析。在ROUGE与F1评分上的评估表明,该级联架构在存在转录错误的情况下仍优于传统摘要方法。未来可优化模型微调与实时处理能力。本研究推动了多模态摘要在信息检索、可访问性与用户体验方面的应用。
原文摘要 · Abstract (English)
This study presents a cascaded architecture for extractive summarization of multimedia content via audio-to-text alignment. The proposed framework addresses the challenge of extracting key insights from multimedia sources like YouTube videos. It integrates audio-to-text conversion using Microsoft Azure Speech with advanced extractive summarization models, including Whisper, Pegasus, and Facebook BART XSum. The system employs tools such as Pytube, Pydub, and SpeechRecognition for content retrieval, audio extraction, and transcription. Linguistic analysis is enhanced through named entity recognition and semantic role labeling. Evaluation using ROUGE and F1 scores demonstrates that the cascaded architecture outperforms conventional summarization methods, despite challenges like transcription errors. Future improvements may include model fine-tuning and real-time processing. This study contributes to multimedia summarization by improving information retrieval, accessibility, and user experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。