让AI自省推理,提升幽默梗图理解能力
Yes FLoReNce, I Will Do Better Next Time! Agentic Feedback Reasoning for Humorous Meme Detection
- 构建闭环反馈机制,让AI能自我修正推理过程
- 在PrideMM数据集上准确率优于静态基线模型
- 无需微调即可通过历史经验优化推理逻辑
幽默梗图融合视觉与文本线索传递讽刺、反讽或社会评论,对AI系统理解意图而非表面关联提出挑战。现有多模态或提示模型虽可生成解释,但推理为开环模式,预测后无法自我批判或优化。本文提出FLoReNce框架,将梗图理解设计为学习阶段的闭环过程、推理阶段的开环过程。在闭环比对中,推理代理由裁判评判,错误与语义反馈转化为控制信号并存入非参数化知识库。推理时,模型从该知识库检索相似判例,用以调制提示,实现无需微调的自对齐推理。在PrideMM数据集上,FLoReNce在预测性能和解释质量上均优于静态多模态基线,表明反馈调控提示是实现自适应梗图幽默理解的有效路径。
原文摘要 · Abstract (English)
Humorous memes blend visual and textual cues to convey irony, satire, or social commentary, posing unique challenges for AI systems that must interpret intent rather than surface correlations. Existing multimodal or prompting-based models generate explanations for humor but operate in an open loop,lacking the ability to critique or refine their reasoning once a prediction is made. We propose FLoReNce, an agentic feedback reasoning framework that treats meme understanding as a closed-loop process during learning and an open-loop process during inference. In the closed loop, a reasoning agent is critiqued by a judge; the error and semantic feedback are converted into control signals and stored in a feedback-informed, non-parametric knowledge base. At inference, the model retrieves similar judged experiences from this KB and uses them to modulate its prompt, enabling better, self-aligned reasoning without finetuning. On the PrideMM dataset, FLoReNce improves both predictive performance and explanation quality over static multimodal baselines, showing that feedback-regulated prompting is a viable path to adaptive meme humor understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。