arXiv:2601.19709cs.SDcs.AI2026-01中稿 · ICASSP 2026被引 1

用双曲空间提升说话人验证的层次特征建模能力

Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification

  • 将说话人嵌入和中心点投影到双曲空间,用双曲距离建模层次结构
  • 双曲加边距Softmax使类间分离度提升,相对EER降低14.23%
  • 适合需要精细层次分类的说话人识别任务

基于欧几里得空间的说话人嵌入学习已取得显著进展,但在建模说话人特征中的层次信息方面仍显不足。双曲空间因其负曲率几何特性,能在有限体积内高效表示层次结构,更适合作为说话人嵌入的特征分布空间。本文提出基于双曲空间的双曲Softmax(H-Softmax)与双曲加边距Softmax(HAM-Softmax)。H-Softmax通过将嵌入与说话人中心投影至双曲空间并计算双曲距离,融入层次信息;HAM-Softmax在此基础上引入边距约束,进一步增强类间可分性。实验表明,H-Softmax与HAM-Softmax相比标准Softmax和AM-Softmax,平均相对EER分别降低27.84%和14.23%,验证了所提方法在提升说话人验证性能的同时,有效保留层次结构建模能力。代码将发布于https://github.com/PunkMale/HAM-Softmax。

原文摘要 · Abstract (English)

Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvature geometric properties, can efficiently represent hierarchical information within a finite volume, making it more suitable for the feature distribution of speaker embeddings. In this paper, we propose Hyperbolic Softmax (H-Softmax) and Hyperbolic Additive Margin Softmax (HAM-Softmax) based on hyperbolic space. H-Softmax incorporates hierarchical information into speaker embeddings by projecting embeddings and speaker centers into hyperbolic space and computing hyperbolic distances. HAM-Softmax further enhances inter-class separability by introducing margin constraint on this basis. Experimental results show that H-Softmax and HAM-Softmax achieve average relative EER reductions of 27.84% and 14.23% compared with standard Softmax and AM-Softmax, respectively, demonstrating that the proposed methods effectively improve speaker verification performance and at the same time preserve the capability of hierarchical structure modeling. The code will be released at https://github.com/PunkMale/HAM-Softmax.

说话人验证双曲空间嵌入学习层次结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。