arXiv:2506.11350cs.SDcs.CL2025-06被引 15

GLAP让音频与文本跨语言跨领域对齐,支持50种语言的语音检索。

GLAP: General contrastive audio-text pretraining across domains and languages

  • 扩展CLAP模型,支持多语言多领域音频-文本预训练。
  • 在50种语言的关键词检测中表现优异,超越现有方法。
  • 适合需要跨语言音频理解的研究者和开发者使用。

对比语言音频预训练(CLAP)是连接音频与文本领域的重要方法。现有CLAP方法主要支持英语中的声音与音乐检索,忽略了多语言语音内容。为此,我们提出通用语言音频预训练(GLAP),扩展了CLAP的多语言与多领域能力。GLAP在标准音频-文本检索基准(如Clotho和AudioCaps)上表现竞争力,同时在语音检索与分类任务中显著优于现有方法。此外,在常用声事件零样本基准上取得强结果,且在语音内容基准上持续领先。跨50种语言的关键词检测评估进一步验证了其先进多语言能力。最后,在四种语言上评估了多语言声音与音乐理解能力。代码与模型权重已公开:https://github.com/xiaomi-research/dasheng-glap。

原文摘要 · Abstract (English)

Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken content. To address this, we introduce general language audio pretraining (GLAP), which expands CLAP with multilingual and multi-domain abilities. GLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and classification tasks. Additionally, GLAP achieves strong results on widely used sound-event zero-shot benchmarks, while simultaneously outperforming previous methods on speech content benchmarks. Further keyword spotting evaluations across 50 languages emphasize GLAP's advanced multilingual capabilities. Finally, multilingual sound and music understanding is evaluated across four languages. Checkpoints and Source: https://github.com/xiaomi-research/dasheng-glap.

多语言音频理解预训练跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。