arXiv:2409.09289cs.SDcs.MM2024-09被引 1

仅用原始音频实现语音文本预训练,提升车载场景智能助手表现。

DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training

  • 通过ASR将音频转为文本,结合对比学习与匹配目标对齐模态。
  • 在1.2万小时车载音频上预训练,下游任务性能全面超越基线。
  • 适合缺乏成对语料的垂直领域语音助手开发,实用性强。

分析真实世界的多模态信号是智能语音助手(IVAs)的关键挑战。主流方法虽在多种下游任务中表现优异,但音视频模型通常独立预训练且基于非目标领域的任务,导致下游任务表征不佳。此外,在许多场景中,收集足够的语言-音频配对数据极为困难,且原始音频转录需高专业技能,使联合预训练难以实现。为此,我们提出DSCLAP,一种仅需原始音频输入即可实现语言-音频预训练的简单高效框架。具体而言,DSCLAP通过ASR系统将原始音频转换为文本,并结合对比学习目标与语言-音频匹配目标,实现音文对齐。我们在12,107小时车载领域音频上进行预训练。实验证明,尽管概念简单,DSCLAP在两个下游任务中的所有指标均显著优于基线模型,展现出在特定领域智能语音助手应用中的巨大潜力。

原文摘要 · Abstract (English)

Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on various downstream tasks of IVAs with pre-trained audio models and text models. However, these models are pre-trained independently and usually on tasks different from target domains, resulting in sub-optimal modality representations for downstream tasks. Moreover, in many domains, collecting enough language-audio pairs is extremely hard, and transcribing raw audio also requires high professional skills, making it difficult or even infeasible to joint pre-training. To address these painpoints, we propose DSCLAP, a simple and effective framework that enables language-audio pre-training with only raw audio signal input. Specifically, DSCLAP converts raw audio signals into text via an ASR system and combines a contrastive learning objective and a language-audio matching objective to align the audio and ASR transcriptions. We pre-train DSCLAP on 12,107 hours of in-vehicle domain audio. Empirical results on two downstream tasks show that while conceptually simple, DSCLAP significantly outperforms the baseline models in all metrics, showing great promise for domain-specific IVAs applications.

语音预训练多模态对齐车载AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。