提出新语音检索任务与关系增强模型,提升情感语调描述的跨模态匹配效果。
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
- 引入关系增强机制,学习语音与描述间的局部匹配关系
- 在情感语调检索任务上超越传统方法,提升跨模态理解能力
- 适合语音生成、情感分析与跨模态检索方向的研究者
对比语言-音频预训练(CLAP)模型在通用音频描述任务中表现优异,但在新兴的情感语调描述(ESSD)领域,跨模态对比预训练仍处于空白。本文提出一种新的语音检索任务——情感语调检索(ESSR),并设计专用于学习语音与自然语言描述间关系的 ESS-CLAP 模型。进一步提出关系增强的 CLAP(RA-CLAP),解决传统方法假设图文间存在严格二元关系的问题。该模型利用自蒸馏机制学习语音与描述之间的潜在局部匹配关系,从而增强泛化能力。实验验证了 RA-CLAP 的有效性,为 ESSD 领域提供了重要参考。
原文摘要 · Abstract (English)
The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。