arXiv:2604.08879cs.CL2026-04

提出GRASP框架,用显式推理定位多模态讽刺目标

GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification

论文配图:GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
图 1 · 摘自论文原文
  • 将视觉定位与思维链结合,显式追踪讽刺推理过程
  • 在MSTI-MAX数据集上显著提升细粒度目标识别准确率
  • 适合研究讽刺识别、可解释AI的学者与工程师

超越传统二分类的多模态讽刺检测,多模态讽刺目标识别(MSTI)需精确定位文本短语和视觉区域等细粒度目标。现有方法主要依赖隐式跨模态对齐,可解释性差且定位不精准。为此,我们提出GRASP:基于双阶段优化的视觉锚定思维链推理框架。首先构建了改进的MSTI-MAX数据集,缓解类别不平衡并丰富多模态讽刺线索;引入视觉锚定的思维链(CoT)推理,明确将视觉区域纳入推理路径,并引导模型在预测前生成理由;采用双阶段监督优化策略:先使用坐标感知加权损失进行微调,再进行细粒度目标策略优化。大量实验表明,GRASP在跨模态细粒度讽刺目标识别上优于现有基线,且通过大模型评判量化了内部推理质量。代码与数据集将在GitHub公开。

原文摘要 · Abstract (English)

Moving beyond the traditional binary classification paradigm of Multimodal Sarcasm Detection, Multimodal Sarcasm Target Identification (MSTI) presents a more formidable challenge, requiring precise localization of fine-grained targets such as textual phrases and visual regions. Existing approaches predominantly rely on implicit cross-modal alignment, offering limited interpretability and suboptimal fine-grained localization. To address these limitations, we propose GRASP, Grounded Chain-of-Thought ReAsoning with Dual-Stage Optimization for Multimodal Sarcasm Prediction and Target Identification, a framework that integrates visual grounding with explicit Chain-of-Thought (CoT) reasoning to move beyond black-box MSTI. Specifically, we curate MSTI-MAX, a refined dataset that mitigates class imbalance and enriches multimodal sarcasm cues. We introduce Grounded CoT reasoning, which explicitly anchors sarcasm-related visual regions within the reasoning trajectory and prompts the model to articulate rationales before predicting the final classification labels and sarcasm targets. Furthermore, we employ a dual-stage outcome-supervised joint optimization strategy: Supervised Fine-Tuning with a coordinate-aware weighted loss, followed by Fine-Grained Target Policy Optimization. Extensive experiments demonstrate that GRASP outperforms existing baselines in fine-grained sarcasm target identification across modalities, and an LLM-as-a-Judge evaluation quantitatively measures the quality of internal reasoning chains. Our dataset and source code will be released on GitHub.

讽刺识别思维链视觉定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。