arXiv:2509.12093cs.CL2025-09中稿 · IEEE ASRU 2025被引 6

开源多语言多模态语义模型,提升语音与文本对齐效果

SENSE models: an open source solution for multilingual and multimodal semantic-based tasks

  • 采用师生框架对齐语音与文本的语义表示
  • 在多语言多模态任务中表现媲美顶尖模型
  • 适合语音识别、跨模态检索等研究者使用

本文提出SENSE(Shared Embedding for N-lingual Speech and tExt),一个受SAMU-XLSR框架启发的开源多语言多模态语义模型,其设计理念与Meta AI的SONAR模型类似。该方法通过师生架构,在话语层面将自监督语音编码器与语言无关的文本编码器连续表示对齐。我们改进了原始SAMU-XLSR方法,选用更强的教师文本模型和更优的初始语音编码器。训练与推理代码已集成至SpeechBrain工具包,首个SENSE模型已公开发布。实验表明,该模型在多语言多模态语义任务中达到极具竞争力的性能。本研究还揭示了此类语义对齐语音编码器中语义表征的捕捉机制。

原文摘要 · Abstract (English)

This paper introduces SENSE (Shared Embedding for N-lingual Speech and tExt), an open-source solution inspired by the SAMU-XLSR framework and conceptually similar to Meta AI's SONAR models. These approaches rely on a teacher-student framework to align a self-supervised speech encoder with the language-agnostic continuous representations of a text encoder at the utterance level. We describe how the original SAMU-XLSR method has been updated by selecting a stronger teacher text model and a better initial speech encoder. The source code for training and using SENSE models has been integrated into the SpeechBrain toolkit, and the first SENSE model we trained has been publicly released. We report experimental results on multilingual and multimodal semantic tasks, where our SENSE model achieves highly competitive performance. Finally, this study offers new insights into how semantics are captured in such semantically aligned speech encoders.

多模态语音编码语义对齐开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。