让大模型理解语音中的情绪和语境,提升共情能力。
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
- 用情绪标注数据显式/隐式训练模型,融合语音语义与情感信息
- 隐式方法使模型问答准确率提升38.41%,联合使用达46.02%
- 适用于需要高共情能力的对话系统与语音助手研发
当前大型语音语言模型在共情推理方面存在局限,主要源于缺乏同时包含上下文内容与副语言线索的训练数据。本文提出两种融入上下文副语言信息的方法:(1) 显式方法,将情绪标注等副语言元数据直接输入大模型;(2) 隐式方法,利用类别与维度化情绪标注及语音转录自动生成新的问答对。隐式方法在人工标注的问答基准上使模型评分提升38.41%,结合显式方法后达到46.02%,证明其在上下文副语言理解上的有效性。我们还通过验证模型评分与分类指标的相关性,支持了评判者可靠性。
原文摘要 · Abstract (English)
Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。