arXiv:2502.00127cs.CL2025-02被引 3

用稀疏自编码器挖掘语音嵌入中的单一语义特征,揭示了语言和音乐等隐藏信息。

Sparse Autoencoder Insights on Voice Embeddings

  • 将稀疏自编码器应用于泰坦网语音嵌入,提取单一语义特征。
  • 发现提取的特征具备特征拆分与可操控性,类似大模型嵌入结果。
  • 适用于语音识别等非文本领域,助力理解嵌入数据内在结构。

近年来可解释机器学习的发展表明,稀疏自编码器在发现密集嵌入中的单语义特征方面具有潜力。尽管多数研究集中于大语言模型(LLM)嵌入,该技术在其他领域的适用性仍待探索。本研究将稀疏自编码器应用于泰坦网(Titanet)生成的说话人嵌入,证明其在非文本嵌入数据中提取单语义特征的有效性。结果表明,所提取特征具备与LLM嵌入相似的特性,包括特征拆分与可操控性。分析显示,自编码器能够识别并操纵原始嵌入中不明显的语言和音乐等特征。研究提示,稀疏自编码器可成为理解与解释多领域嵌入数据的有力工具,尤其在基于音频的说话人识别中具有应用前景。

原文摘要 · Abstract (English)

Recent advances in explainable machine learning have highlighted the potential of sparse autoencoders in uncovering mono-semantic features in densely encoded embeddings. While most research has focused on Large Language Model (LLM) embeddings, the applicability of this technique to other domains remains largely unexplored. This study applies sparse autoencoders to speaker embeddings generated from a Titanet model, demonstrating the effectiveness of this technique in extracting mono-semantic features from non-textual embedded data. The results show that the extracted features exhibit characteristics similar to those found in LLM embeddings, including feature splitting and steering. The analysis reveals that the autoencoder can identify and manipulate features such as language and music, which are not evident in the original embedding. The findings suggest that sparse autoencoders can be a valuable tool for understanding and interpreting embedded data in many domains, including audio-based speaker recognition.

语音嵌入稀疏自编码器可解释性特征挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。