让语音模型能生成带方言、年龄等细节的描述性画像
CoLMbo: Speaker Language Model for Descriptive Profiling
- 用提示词动态适配,将语音嵌入转化为结构化描述
- 零样本下跨数据集生成方言与年龄特征描述
- 适合需要个性化语音画像的场景,如智能客服
语音识别系统通常仅限于分类任务,难以生成详细的说话人特征或提供上下文丰富的描述。现有模型主要提取用于身份识别的嵌入,却无法以结构化方式捕捉性别、年龄、方言等人口统计属性。本文提出CoLMbo,一种说话人语言模型(SLM),通过将说话人编码器与基于提示的条件机制结合,实现从说话人嵌入生成详细描述。用户可自定义提示,使模型动态适应新特征,生成包含地域方言差异和年龄相关特征的定制化描述。该方法不仅提升传统说话人画像能力,更在多个数据集上展现出优异的零样本表现,为说话人识别领域带来重要进展。
原文摘要 · Abstract (English)
Speaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models primarily extract embeddings for speaker identification but fail to capture demographic attributes such as dialect, gender, and age in a structured manner. This paper introduces CoLMbo, a Speaker Language Model (SLM) that addresses these limitations by integrating a speaker encoder with prompt-based conditioning. This allows for the creation of detailed captions based on speaker embeddings. CoLMbo utilizes user-defined prompts to adapt dynamically to new speaker characteristics and provides customized descriptions, including regional dialect variations and age-related traits. This innovative approach not only enhances traditional speaker profiling but also excels in zero-shot scenarios across diverse datasets, marking a significant advancement in the field of speaker recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。