用关键词增强大模型嵌入,实现开放集说话人属性预测
Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings
- 用大语言模型嵌入表示属性,构建连续语义空间
- 引入关键词附加策略,提升嵌入的区分度与紧凑性
- 在LibriTTS-P上对未见同义词有效泛化,适合开放场景
理解说话人属性对语音应用至关重要,但传统方法依赖固定类别标签,缺乏语义丰富性与零样本泛化能力。本文提出一种新型开放集说话人属性预测框架,利用大语言模型(LLM)嵌入将属性映射到连续语义空间。为弥合跨模态差异,引入关键词附加策略,将宽泛语义表示压缩为紧凑、可区分的流形结构。同时采用top-k负损失,在密集语义区域建立稳健决策边界。在LibriTTS-P上的实验表明,该方法超越封闭集基准,并有效泛化至未见同义词。几何分析显示,所提策略正则化嵌入流形,平衡语义凝聚性与预测清晰性。
原文摘要 · Abstract (English)
Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot generalizability. We propose a novel framework for open-set speaker attribute prediction leveraging Large Language Model (LLM) embeddings to represent attributes in a continuous semantic space. To bridge the cross-modal gap, we introduce a keyword-appending strategy that structures broad semantic representations into a compact, discriminative manifold. Furthermore, we employ a top-k negative loss to establish robust decision boundaries in crowded semantic regions. Experimental results on LibriTTS-P demonstrate that our method outperforms closed-set benchmarks and generalizes effectively to unseen synonyms. Geometric analysis suggests that our strategies regularize the embedding manifold, balancing semantic cohesion with predictive clarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。