用大模型对齐人类偏好,让语音情感描述更真实准确
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
- 用知识蒸馏对齐语音与文本的语义分布
- 通过偏好优化减少幻觉,提升描述真实性
- 适合需要高可信度语音情感分析的研究者
语音情感描述(SEC)正成为研究热点。人类语音传递的情感通常复杂,固定分类难以充分捕捉。用自然语言描述情感更具优势。但现有方法常出现幻觉且在未见语音上泛化能力差。为此,我们提出AlignCap,基于大语言模型实现语音情感描述对齐人类偏好,具备两个特性:1)语音-文本对齐,通过知识蒸馏正则化最小化大模型对语音和文本输入的响应分布差异;2)人类偏好对齐,设计偏好优化正则化以消除事实性与忠实性幻觉。同时,提取情感线索作为提示,在知识蒸馏正则化下丰富细粒度信息。实验表明,AlignCap在零样本语音情感描述任务上优于现有最先进方法。
原文摘要 · Abstract (English)
Speech Emotion Captioning (SEC) has gradually become an active research task. The emotional content conveyed through human speech are often complex, and classifying them into fixed categories may not be enough to fully capture speech emotions. Describing speech emotions through natural language may be a more effective approach. However, existing SEC methods often produce hallucinations and lose generalization on unseen speech. To overcome these problems, we propose AlignCap, which Aligning Speech Emotion Captioning to Human Preferences based on large language model (LLM) with two properties: 1) Speech-Text Alignment, which minimizing the divergence between the LLM's response prediction distributions for speech and text inputs using knowledge distillation (KD) Regularization. 2) Human Preference Alignment, where we design Preference Optimization (PO) Regularization to eliminate factuality and faithfulness hallucinations. We also extract emotional clues as a prompt for enriching fine-grained information under KD-Regularization. Experiments demonstrate that AlignCap presents stronger performance to other state-of-the-art methods on Zero-shot SEC task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。