arXiv:2410.12948cs.CLcs.SD2024-10被引 9

分析语音大模型没学懂的非语言信息,揭示其表征能力与适配潜力。

What Do Speech Foundation Models Not Learn About Speech?

  • 通过动态超级基准测试,评估多模型在零样本下的非语言线索捕捉能力。
  • 发现模型深层特征呈现凸性分离规律,不同层捕获任务特异性表征。
  • 零样本表现好预示表征质量高,且可有效微调适配下游任务。

理解语音基础模型如何捕捉非语言线索对提升其可解释性与跨任务适应性至关重要。本文分析Whisper、Seamless、Wav2Vec、HuBERT和Qwen2-Audio等主流模型,聚焦其在Dynamic-SUPERB基准下对语气、情感、环境上下文等非语言任务的表征能力。研究回答三个问题:(1) 模型捕捉哪些非语言线索?(2) 这些线索在模型各层如何分布?(3) 表征能否有效适配下游任务?我们先在零样本设置下评估模型,再对各层特征进行微调。结果表明,尽管未显式训练于这些任务,部分模型仍表现出良好零样本性能,且该性能与学习到的表征质量正相关。层间分析显示,某些模型存在表征可分性随深度呈凸变化的现象,不同层提取任务特定特征。

原文摘要 · Abstract (English)

Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze several prominent models such as Whisper, Seamless, Wav2Vec, HuBERT, and Qwen2-Audio focusing on their learned representations in both paralinguistic and non-paralinguistic tasks from the Dynamic-SUPERB benchmark. Our study addresses three key questions: (1) What non-verbal cues (e.g., speaker intent, emotion, environmental context) are captured? (2) How are these cues represented across different layers of the models? and (3) To what extent can these representations be effectively adapted to downstream tasks? To answer these questions, we first evaluate the models in a zero-shot setting, followed by fine-tuning on layer-wise features extracted from these models. Our results provide insights into the models' capacity for generalization, the characteristics of their layer-wise representations, and the degree of transformation required for downstream task adaptation. Our findings suggest that some of these models perform well on various tasks in zero-shot settings, despite not being explicitly trained for those tasks. We also observe that zero-shot performance correlates with better-learned representations. The analysis of layer-wise features demonstrates that some models exhibit a convex relationship between the separability of the learned representations and model depth, with different layers capturing task-specific features.

语音模型非语言线索表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。