梳理语音摘要研究现状,揭示技术演进与核心挑战。
Summarizing Speech: A Comprehensive Survey
- 整合语音识别与文本摘要技术,构建端到端生成框架
- 现有数据集和评估标准仍不完善,制约模型发展
- 适合关注语音处理、多模态分析的研究者阅读
语音摘要已成为高效管理日益增长的语音及音视频内容的重要工具。尽管其重要性不断提升,但该领域定义仍较模糊,涉及语音识别、文本摘要及会议摘要等多重研究方向。本综述不仅分析现有数据集与评估协议,这些对衡量摘要质量至关重要,还总结了近期进展,突出从传统系统向微调级联架构和端到端方案的转变。同时,我们揭示了当前挑战,包括真实评估基准的缺乏、多语言数据集的不足以及长上下文处理难题。
原文摘要 · Abstract (English)
Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization remains loosely defined. The field intersects with several research areas, including speech recognition, text summarization, and specific applications like meeting summarization. This survey not only examines existing datasets and evaluation protocols, which are crucial for assessing the quality of summarization approaches, but also synthesizes recent developments in the field, highlighting the shift from traditional systems to advanced models like fine-tuned cascaded architectures and end-to-end solutions. In doing so, we surface the ongoing challenges, such as the need for realistic evaluation benchmarks, multilingual datasets, and long-context handling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。