arXiv:2603.05275cs.MMcs.CL2026-03

用强化学习提升多模态讽刺识别,防止幻觉并增强推理可信度。

SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning

  • 将讽刺识别重构为结构化推理任务,双轨蒸馏初始化模型
  • 在MUStARD++上F1达70.22%,超越零样本和微调基线
  • 适合需要高可信多模态推理的场景,如智能客服、内容审核

多模态讽刺检测需通过跨模态推理解决文本、声学与视觉线索间的语用不一致问题。为使基础模型实现稳健的讽刺推理,我们提出SarcasmMiner,一种基于强化学习的后训练框架,可抑制多模态推理中的幻觉现象。我们将讽刺检测重构为结构化推理任务,并采用双轨蒸馏策略:高质量教师轨迹初始化学生模型,全量轨迹用于训练生成式奖励模型(GenRM)以评估推理质量。学生模型通过解耦的准确率与推理质量奖励,使用群体相对策略优化(GRPO)进行优化。在MUStARD++数据集上,SarcasmMiner将F1从零样本的59.83%、监督微调的68.23%提升至70.22%。结果表明,面向推理的奖励建模能同时提升性能与多模态对齐性。

原文摘要 · Abstract (English)

Multimodal sarcasm detection requires resolving pragmatic incongruity across textual, acoustic, and visual cues through cross-modal reasoning. To enable robust sarcasm reasoning with foundation models, we propose SarcasmMiner, a reinforcement learning based post-training framework that resists hallucination in multimodal reasoning. We reformulate sarcasm detection as structured reasoning and adopt a dual-track distillation strategy: high-quality teacher trajectories initialize the student model, while the full set of trajectories trains a generative reward model (GenRM) to evaluate reasoning quality. The student is optimized with group relative policy optimization (GRPO) using decoupled rewards for accuracy and reasoning quality. On MUStARD++, SarcasmMiner increases F1 from 59.83% (zero-shot), 68.23% (supervised finetuning) to 70.22%. These findings suggest that reasoning-aware reward modeling enhances both performance and multimodal grounding.

多模态讽刺识别强化学习推理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。