评估多模态大模型对语音讽刺的理解能力,发现音频信息最有效。
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
- 用融合门控模块整合文本、语音、视觉特征,提升讽刺识别效果。
- 语音单模态表现最佳,文本+语音组合优于三模态整体。
- 中文与英文数据集上,多模态大模型在零样本和微调下均表现良好。
讽刺理解仍是自然语言理解的挑战,因其意图常依赖文本、语音和视觉间的细微跨模态线索。现有研究多集中于文本或视觉-文本讽刺,而全面的音视频-文本讽刺理解仍待探索。本文系统评估了大语言模型(LLMs)和多模态大语言模型(MLLMs)在英语(MUStARD++)和中文(MCSD 1.0)数据集上的讽刺检测性能,涵盖零样本、少样本及LoRA微调设置。除直接分类外,还探索将模型作为特征编码器,通过协同门控融合模块整合其表示。实验结果表明,基于音频的模型在单模态中表现最强,文本-音频与音频-视觉组合优于单模态及三模态模型。此外,如Qwen-Omni等多模态大模型在零样本和微调场景下表现具有竞争力。研究结果凸显了多模态大模型在跨语言、音视频-文本讽刺理解中的潜力。
原文摘要 · Abstract (English)
Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual sarcasm, comprehensive audio-visual-textual sarcasm understanding remains underexplored. In this paper, we systematically evaluate large language models (LLMs) and multimodal LLMs for sarcasm detection on English (MUStARD++) and Chinese (MCSD 1.0) in zero-shot, few-shot, and LoRA fine-tuning settings. In addition to direct classification, we explore models as feature encoders, integrating their representations through a collaborative gating fusion module. Experimental results show that audio-based models achieve the strongest unimodal performance, while text-audio and audio-vision combinations outperform unimodal and trimodal models. Furthermore, MLLMs such as Qwen-Omni show competitive zero-shot and fine-tuned performance. Our findings highlight the potential of MLLMs for cross-lingual, audio-visual-textual sarcasm understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。