跨语言区分口语与剧本式语音,提升媒体内容理解与推荐
Classification of Spontaneous and Scripted Speech for Multilingual Audio
- 用音频变换器模型捕捉多语言语音风格差异
- 在11种语言上均达当前最优效果,超越传统声学特征方法
- 适合做语音分析、媒体内容管理的工程师和研究者
区分剧本式语音与即兴口语是理解语音风格对语音处理影响的关键。该任务还能通过更精准分割大型语音数据库,提升推荐系统和媒体发现体验。本文针对跨语言、跨格式的分类器泛化能力挑战,系统评估了从传统手工声学/语调特征到先进音频变换器模型的表现。基于大规模多语言专有播客数据集训练与验证,按11个语言组拆解模型性能,分析跨语言偏差。实验还扩展至公开数据集,检验模型在非播客场景下的泛化能力。结果表明,变换器模型始终优于传统方法,在多种语言中实现顶尖性能。
原文摘要 · Abstract (English)
Distinguishing scripted from spontaneous speech is an essential tool for better understanding how speech styles influence speech processing research. It can also improve recommendation systems and discovery experiences for media users through better segmentation of large recorded speech catalogues. This paper addresses the challenge of building a classifier that generalises well across different formats and languages. We systematically evaluate models ranging from traditional, handcrafted acoustic and prosodic features to advanced audio transformers, utilising a large, multilingual proprietary podcast dataset for training and validation. We break down the performance of each model across 11 language groups to evaluate cross-lingual biases. Our experimental analysis extends to publicly available datasets to assess the models' generalisability to non-podcast domains. Our results indicate that transformer-based models consistently outperform traditional feature-based techniques, achieving state-of-the-art performance in distinguishing between scripted and spontaneous speech across various languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。