让语音模型懂情绪、方言和说话人特征,更像真人对话。
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
- 双模态头设计分离语言与语音表现,提升理解与生成能力。
- 在多维度评测中,对情感、方言、年龄互动表现优于现有开源模型。
- 适合开发更自然、有社交意识的语音交互系统。
端到端语音语言模型虽显著提升了人工智能的自然语音交互能力,但多数模型仅将语音视为语言内容的载体,忽视了语调、方言、年龄、情绪及非语音发声等丰富的副语言与说话人特征。本文提出GOAT-SLM,一种具备副语言与说话人特征感知能力的新型语音语言模型,旨在拓展语音建模超越文本语义的范畴。该模型采用双模态头架构,解耦语言建模与声学实现,支持鲁棒的语言理解与富有表现力的语音生成。通过基于大规模语音-文本语料库的模块化分阶段训练策略,逐步对齐语言、副语言与说话人信息,提升效率与泛化性。在TELEVAL多维评估基准上的实验表明,GOAT-SLM在语义与非语义任务间实现良好平衡,且在情感识别、方言变化、年龄敏感交互等方面优于现有开源模型。本研究强调超越语言内容建模的重要性,推动更具自然性、适应性与社会感知能力的语音系统发展。
原文摘要 · Abstract (English)
Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing models treat speech merely as a vehicle for linguistic content, often overlooking the rich paralinguistic and speaker characteristic cues embedded in human speech, such as dialect, age, emotion, and non-speech vocalizations. In this work, we introduce GOAT-SLM, a novel spoken language model with paralinguistic and speaker characteristic awareness, designed to extend spoken language modeling beyond text semantics. GOAT-SLM adopts a dual-modality head architecture that decouples linguistic modeling from acoustic realization, enabling robust language understanding while supporting expressive and adaptive speech generation. To enhance model efficiency and versatility, we propose a modular, staged training strategy that progressively aligns linguistic, paralinguistic, and speaker characteristic information using large-scale speech-text corpora. Experimental results on TELEVAL, a multi-dimensional evaluation benchmark, demonstrate that GOAT-SLM achieves well-balanced performance across both semantic and non-semantic tasks, and outperforms existing open-source models in handling emotion, dialectal variation, and age-sensitive interactions. This work highlights the importance of modeling beyond linguistic content and advances the development of more natural, adaptive, and socially aware spoken language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。