用自然语言监督训练语音模型,实现零样本跨语言健康与说话人识别。
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
- 通过对比学习对齐语音与文本描述的说话人/健康特征
- 零样本平均F1达62.9%,比CLAP提升48%;临床任务最优57.9% F1
- 适用于跨语言、跨人群的语音健康分析,适合医疗语音研究者
语音包含年龄、性别、声音质量及健康状态等副语言信息。然而,现有音频基础模型无法支持零样本或分布外(OOD)泛化。我们提出SLAP(Speaker contrastive Language-Audio Pretraining),首个通过对比学习将语音与说话人及健康元数据的自然语言描述对齐的模型。SLAP结合视觉变压器音频编码器与文本编码器,在超过3400小时、涵盖9个数据集、具有多样化说话人标注的数据上训练。在14个语言、7个数据集上的38个二分类任务中评估,包含人口统计、语音特征和临床评估。零样本下平均F1为62.9%,相较CLAP(42.4%)提升48%;在未见语言和临床人群中仍表现稳健。线性探测微调后总体F1达69.3%,健康任务最佳表现57.9% F1,超越更大规模基础模型。
原文摘要 · Abstract (English)
Speech encodes paralinguistic information such as demographics, voice quality, and health. Yet no audio foundation model supports zero-shot or out-of-distribution (OOD) generalization to these tasks. We introduce SLAP (Speaker contrastive Language-Audio Pretraining), the first model aligning speech with natural language descriptions of speaker and health metadata through contrastive learning. SLAP combines a Vision Transformer audio encoder with text encoders, trained on more than 3400 hours across 9 datasets with diverse speaker annotations. We evaluated on 38 binary classification tasks spanning demographics, voice characteristics, and clinical assessments across 14 datasets in 7 languages. SLAP achieves 62.9% average F1 in zero-shot evaluation, a 48% relative improvement over CLAP (42.4%), while demonstrating strong OOD generalization to unseen languages and clinical populations. When fine-tuned with linear probing, SLAP reaches 69.3% F1 overall and achieves best-in-class performance on health tasks (57.9% F1), surpassing larger foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。