构建诊断基准,评估视频模型识别讽刺意图的推理能力
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
- 提出可分解感知与推理的诊断框架
- 在多模态讽刺检测上提升模型推理准确率
- 适合研究多模态理解与隐含意图识别的学者
讽刺是一种特定的反语形式,需区分字面意义与实际含义。检测讽刺不仅依赖话语本身,还需结合语气、面部表情和对话上下文等非语言线索。当前多模态模型在复杂任务如讽刺检测中表现不佳,难以跨模态识别相关线索并进行语用推理以判断说话人意图。为此,我们引入MUStReason,一个包含模态特异性线索标注和推理步骤注释的诊断基准,用于分析视频语言模型(VideoLMs)在讽刺检测中的表现。通过该基准,我们定量与定性评估模型生成的推理过程,将问题拆解为感知与推理两部分,并提出PragCoT框架,引导模型关注隐含意图而非字面意义——这是识别讽刺的核心特征。
原文摘要 · Abstract (English)
Sarcasm is a specific type of irony which involves discerning what is said from what is meant. Detecting sarcasm depends not only on the literal content of an utterance but also on non-verbal cues such as speaker's tonality, facial expressions and conversational context. However, current multimodal models struggle with complex tasks like sarcasm detection, which require identifying relevant cues across modalities and pragmatically reasoning over them to infer the speaker's intention. To explore these limitations in VideoLMs, we introduce MUStReason, a diagnostic benchmark enriched with annotations of modality-specific relevant cues and underlying reasoning steps to identify sarcastic intent. In addition to benchmarking sarcasm classification performance in VideoLMs, using MUStReason we quantitatively and qualitatively evaluate the generated reasoning by disentangling the problem into perception and reasoning, we propose PragCoT, a framework that steers VideoLMs to focus on implied intentions over literal meaning, a property core to detecting sarcasm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。