arXiv:2506.21712cs.CLcs.SD2025-06被引 4

发现自监督语音模型中编码说话人信息的关键神经元

Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers

  • 通过聚类自监督特征与i-vectors定位关键神经元
  • 保护这些神经元可显著维持说话人任务性能
  • 适合研究语音表征与模型可解释性的学者

近年来,自监督语音Transformer在说话人相关任务中取得进展,但其如何编码说话人信息仍不明确。本文通过分析自监督特征的k-means聚类与i-vectors对应的神经元,发现这些聚类对应广泛的语音和性别类别,因而可用于识别代表说话人的神经元。在剪枝过程中保护这些神经元,可显著保留说话人相关任务的性能,证明其在编码说话人信息中的关键作用。

原文摘要 · Abstract (English)

In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information.

语音模型神经元分析说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。