arXiv:2510.24178cs.CLcs.AI2025-10被引 1

首个德语多模态反讽数据集,助力模型理解真实语境中的反讽。

MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations

  • 构建包含文本、音频、视频的德语反讽数据集,每条标注独立且对齐。
  • 模型在文本上表现最佳,但人类更依赖音频线索,暴露当前多模态模型短板。
  • 适合研究多模态反讽检测、人机判断差异及跨模态理解的学者使用。

反讽是一种字面意义与实际意图相悖的复杂修辞形式,在社交媒体和流行文化中广泛存在,给自然语言理解、情感分析和内容审核带来持续挑战。随着多模态大模型的发展,反讽检测已超越纯文本,需融合语音与视觉线索。我们提出MuSaG,首个德语多模态反讽检测数据集,包含33分钟来自德国电视节目的手动筛选与人工标注语句。每个实例提供对齐的文本、音频和视频模态,由人类分别标注,支持单模态与多模态评估。我们对九种开源及商用模型(涵盖文本、音频、视觉与多模态架构)进行基准测试,并与人类标注结果对比。结果显示,人类在对话场景中高度依赖音频,而模型在文本任务上表现最优,凸显当前多模态模型的不足,也激励未来模型更贴近真实场景。我们公开发布MuSaG,以推动多模态反讽检测与人-模型对齐研究。

原文摘要 · Abstract (English)

Sarcasm is a complex form of figurative language in which the intended meaning contradicts the literal one. Its prevalence in social media and popular culture poses persistent challenges for natural language understanding, sentiment analysis, and content moderation. With the emergence of multimodal large language models, sarcasm detection extends beyond text and requires integrating cues from audio and vision. We present MuSaG, the first German multimodal sarcasm detection dataset, consisting of 33 minutes of manually selected and human-annotated statements from German television shows. Each instance provides aligned text, audio, and video modalities, annotated separately by humans, enabling evaluation in unimodal and multimodal settings. We benchmark nine open-source and commercial models, spanning text, audio, vision, and multimodal architectures, and compare their performance to human annotations. Our results show that while humans rely heavily on audio in conversational settings, models perform best on text. This highlights a gap in current multimodal models and motivates the use of MuSaG for developing models better suited to realistic scenarios. We release MuSaG publicly to support future research on multimodal sarcasm detection and human-model alignment.

反讽检测多模态德语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。