用大模型融合语音转写上下文与多系统输出,显著提升情感识别准确率。
Context and System Fusion in Post-ASR Emotion Recognition with Large Language Models
- 通过排序转写文本和动态调整对话上下文,优化输入提示。
- 在GenSEC任务上,最佳方案比基线高出20%绝对准确率。
- 适合做语音情感分析、多模态交互系统研发的团队参考。
大型语言模型(LLMs)在语音与文本建模中扮演着日益重要的角色。为探索上下文与多个系统输出在后语音识别(post-ASR)情感预测中的最佳利用方式,我们研究了在近期任务GenSEC上使用LLM提示的方法。技术包括语音转写文本排序、可变对话上下文构建以及系统输出融合。结果表明,对话上下文存在边际效益递减现象,且用于选择预测转写的评估指标至关重要。最终,我们的最优提交方案在绝对准确率上超越提供基线20%。
原文摘要 · Abstract (English)
Large language models (LLMs) have started to play a vital role in modelling speech and text. To explore the best use of context and multiple systems' outputs for post-ASR speech emotion prediction, we study LLM prompting on a recent task named GenSEC. Our techniques include ASR transcript ranking, variable conversation context, and system output fusion. We show that the conversation context has diminishing returns and the metric used to select the transcript for prediction is crucial. Finally, our best submission surpasses the provided baseline by 20% in absolute accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。