arXiv:2606.03604cs.CL2026-06

让AI看懂梗图背后的潜台词,而非只描述画面。

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

论文配图:Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding
图 1 · 摘自论文原文
  • 用正交投影分离图像文字表层与隐含意图信号。
  • 在六项基准测试中超越开源模型,高分歧梗图效果提升显著。
  • 适合研究幽默理解、多模态推理或对齐人类语义的学者。

当被问及梗图或讽刺内容的含义时,大型视觉语言模型往往仅描述图像表面内容,而忽略作者的真实意图。标准指令微调将字面信息与语用意义混淆,导致表层细节干扰最终回答。本文将梗图理解重构为字面-语用分解问题,提出「意图投射」框架,在单一LVLM主干中从表示、输出到目标三个层面实现分离。表示层通过正交投影模块去除融合表征中的主导单模态方向,仅保留语用残差;同时以表面真实情感分类器锚定解码器,输出极性差异标签。输出层外化结构化推理链,目标层采用对比奖励显式惩罚重复字面描述的答案。在六个多模态基准上,该方法持续优于开源基线,并缩小与专有模型的差距,尤其在高分歧帖子上提升最显著。

原文摘要 · Abstract (English)

When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Standard instruction tuning entangles a post's literal content with its pragmatic meaning, letting surface-level details contaminate the final response. We reframe meme understanding as a problem of literal-pragmatic decomposition and propose \textbf{Intent Projection}, a framework that separates the two signals at the representation, output, and objective levels within a single LVLM backbone. At the representation level, an orthogonal projection module removes dominant unimodal directions from the fused image-text representation, retaining only the pragmatic residual, while a surface-real affect classifier anchors the decoder with a discrete tag that names the polarity gap. At the output level, the model externalizes a structured reasoning chain, and at the objective level a contrastive reward explicitly penalizes answers that restate the literal description. Across six multimodal benchmarks, Intent Projection consistently outperforms open-source baselines and narrows the gap to proprietary models, with the largest gains on high-divergence posts where literal collapse is most damaging.

多模态理解语用分析梗图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。