打造能同时理解语音身份与语气的通用声纹编码器
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
- 多任务训练让声纹表示更均衡,兼顾身份与语调信息
- 对比语言-音频预训练仅提升检索能力,未改善语气理解
- 可无缝接入大语言模型,适合语音理解研究者使用
人类语音同时包含身份和副语言特征,但现有大型音频-语言模型中的编码器通常难以平衡这两方面。本文研究构建通用声纹编码器以捕捉细微语音线索。通过全面评估发现,多任务训练能实现最均衡的表征,而对比语言-音频预训练(CLAP)虽提升检索性能,却未增强副语言理解能力。最终提出的Auden-Voice编码器在集成至大语言模型时表现优异。代码与训练方法将随音频理解工具包Auden一同发布。
原文摘要 · Abstract (English)
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。