arXiv:2503.21480cs.CL2025-03被引 9

用通用大模型零样本识别情绪,效果媲美微调音频模型。

OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

  • 用通用大模型处理语音情绪,无需微调即可跨模态推理。
  • 在IEMOCAP和MELD数据集上,零样本性能媲美甚至超越微调音频模型。
  • 提出声学提示策略,提升模型对语音特征与对话上下文的理解。

目前对通用大语言模型(omni-LLMs)在多模态认知状态任务中的应用研究不足,尤其在语音情绪识别方面。本文首次系统评估了四种 omni-LLMs 在零样本情绪识别任务上的表现。在 IEMOCAP 和 MELD 两个广泛使用的多模态情绪基准上,零样本 omni-LLMs 的表现优于或不逊于微调的音频模型。除仅用语音的评估外,还考察了纯文本及文本+语音组合场景。提出声学提示(acoustic prompting),聚焦声学特征分析、对话上下文理解与分步推理。对比最小提示与完整思维链提示,发现声学提示更有效。在 IEMOCAP 与 MELD 上进行上下文窗口分析,表明使用上下文有助于提升性能,尤其在 IEMOCAP 上效果显著。最后对 omni-LLMs 生成的声学推理输出进行误差分析。

原文摘要 · Abstract (English)

The use of omni-LLMs (large language models that accept any modality as input), particularly for multimodal cognitive state tasks involving speech, is understudied. We present OmniVox, the first systematic evaluation of four omni-LLMs on the zero-shot emotion recognition task. We evaluate on two widely used multimodal emotion benchmarks: IEMOCAP and MELD, and find zero-shot omni-LLMs outperform or are competitive with fine-tuned audio models. Alongside our audio-only evaluation, we also evaluate omni-LLMs on text only and text and audio. We present acoustic prompting, an audio-specific prompting strategy for omni-LLMs which focuses on acoustic feature analysis, conversation context analysis, and step-by-step reasoning. We compare our acoustic prompting to minimal prompting and full chain-of-thought prompting techniques. We perform a context window analysis on IEMOCAP and MELD, and find that using context helps, especially on IEMOCAP. We conclude with an error analysis on the generated acoustic reasoning outputs from the omni-LLMs.

情绪识别通用大模型零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。