arXiv:2412.13071cs.CLcs.IR2024-12中稿 · ECIR 2025, 13 page…被引 3

CLASP让语音和文本跨语言检索更准,不依赖语音转文字。

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

  • 用语音频谱+预训练语音模型融合文本编码,统一处理多语言多模态数据。
  • 在15类语料上测试,多项指标超越传统语音转写方法,尤其在低资源场景更优。
  • 适合做跨语言语音搜索、多模态信息检索的研究者与工程师使用。

本文提出CLASP(对比语言-语音预训练),一种面向多语言多模态信息检索的表示学习模型。该模型利用语音内容与文本数据的协同关系,在包含15个类别(如小说、宗教)的新建语音-文本数据集上进行训练。其音频部分结合音频频谱图与预训练自监督语音模型,文本部分采用在超过100种语言上预训练的句子编码器。该轻量级统一模型有效弥合了不同模态与语言间的鸿沟,显著提升多语言多模态数据的检索能力。在多种语言上的评估显示,CLASP在HITS@1、MRR和meanR等指标上建立新基准,优于依赖语音转写后进行文本检索的传统方法,尤其在特定场景下表现突出。

原文摘要 · Abstract (English)

This study introduces CLASP (Contrastive Language-Speech Pretraining), a multilingual, multimodal representation tailored for audio-text information retrieval. CLASP leverages the synergy between spoken content and textual data. During training, we utilize our newly introduced speech-text dataset, which encompasses 15 diverse categories ranging from fiction to religion. CLASP's audio component integrates audio spectrograms with a pre-trained self-supervised speech model, while its language encoding counterpart employs a sentence encoder pre-trained on over 100 languages. This unified lightweight model bridges the gap between various modalities and languages, enhancing its effectiveness in handling and retrieving multilingual and multimodal data. Our evaluations across multiple languages demonstrate that CLASP establishes new benchmarks in HITS@1, MRR, and meanR metrics, outperforming traditional ASR-based retrieval methods that rely on transcribing speech into text for subsequent text retrieval, especially in specific scenarios.

多模态跨语言语音检索预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。