用自然语言描述风格,精准检索特定表达方式的语音片段。
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
- 构建语音与文本编码器,将语音和风格描述映射到共享语义空间。
- 在22种说话风格上实现高召回率,Recall@k表现优异。
- 适合需要按情感或语调检索语音的研究者与开发者。
我们提出表达性语音检索任务:根据对说话风格的自然语言描述,检索出具有该风格的语音片段。以往工作主要基于语音内容进行检索,而本研究关注的是语音的表达方式。通过训练语音和文本编码器,将语音及风格描述嵌入联合隐空间,使自由形式的文本提示(如描述情绪或语调)可作为查询来匹配对应语音段。我们在多个包含22种说话风格的数据集上进行了详尽分析,涵盖编码器架构、跨模态对齐训练策略以及提示增强方法,以提升对任意文本查询的泛化能力。实验表明,该方法在召回率@k指标上表现强劲。
原文摘要 · Abstract (English)
We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused on performing speech retrieval based on what was said in an utterance, we aim to do so based on how something was said. We train speech and text encoders to embed speech and text descriptions of speaking styles into a joint latent space, which enables using free-form text prompts describing emotions or styles as queries to retrieve matching expressive speech segments. We perform detailed analyses of various aspects of our proposed framework, including encoder architectures, training criteria for effective cross-modal alignment, and prompt augmentation for improved generalization to arbitrary text queries. Experiments on multiple datasets encompassing 22 speaking styles demonstrate that our approach achieves strong retrieval performance as measured by Recall@k.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。