让模型自动生成分析视角,动态修正推理过程,提升多模态讽刺检测准确率。
ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection

- 模型自主生成多角度分析视角,取代固定预设模板。
- 通过批评者反馈实现推理过程定向修正,提升判断可靠性。
- 适用于需要深度跨模态理解的讽刺内容识别任务。
多模态讽刺检测需分析字面表达与真实意图之间的跨模态矛盾,但讽刺机制多样,所需分析视角各不相同。现有方法虽显式化分析过程,仍依赖固定预设视角和人工设计的路由规则。本文提出ProCrit,一种提案-批评者双代理框架:提案代理生成多视角推理,批评者提供外部评估与定向修订建议。为解决现有数据集缺乏过程级标注的问题,ProCrit通过动态角色代理演进机制合成推理标注——强视觉语言模型在共享上下文中逐次生成分析角色,形成保留跨视角依赖关系的序列,支持高效自回归生成。为提高推理可靠性,采用草稿-批评-修订范式,独立批评者识别推理缺陷并提供自然语言反馈以指导修订。最后,设计互优化训练框架,通过双阶段强化学习联合优化提案与反馈引导的修订,并根据反馈实际效果持续优化批评者。在三个主流基准上的实验验证了ProCrit的有效性。
原文摘要 · Abstract (English)
Multimodal sarcasm detection requires reasoning over cross-modal incongruities between literal expression and intended meaning, yet the specific analytical perspectives needed vary across samples due to the diversity of sarcastic mechanisms. While recent methods make this analytical process explicit, they still rely on fixed, predefined perspectives that operate independently under hand-crafted routing rules. We argue that multimodal sarcasm detection instead calls for self-elicited multi-perspective reasoning, where a model autonomously generates the perspectives needed for each sample and progressively integrates them into a coherent analysis. To realize this goal, we propose ProCrit, a Proposal-Critic two-agent framework with a proposal agent for multi-perspective reasoning and a critic agent for external evaluation and targeted revision guidance. First, to overcome the lack of process-level supervision in existing sarcasm datasets, ProCrit synthesizes process-level reasoning annotations through a dynamic-role agentic rollout: a strong vision-language model sequentially spawns analytical roles within a shared context, and the resulting multi-role trajectories are flattened into sequences that preserve cross-perspective dependencies while enabling efficient autoregressive generation. Second, to improve reasoning reliability, ProCrit adopts a draft-critique-revise paradigm in which an independent critic identifies reasoning deficiencies and provides targeted natural-language feedback for directed revision. Finally, we develop a mutual-refinement training framework that jointly optimizes proposal drafting and feedback-guided revision via dual-stage reinforcement learning, while refining the critic agent according to the actual effectiveness of its feedback. Experiments on three widely used benchmarks demonstrate the effectiveness of ProCrit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。