arXiv:2601.17197cs.CLcs.LG2026-01Conference of the …

让AI读懂图文中的隐喻讽刺,还能解释推理过程。

Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding

  • 设计三步框架,跨模态理解隐喻、幽默等抽象表达
  • 加入推理路径后,多风格理解能力显著提升
  • 小模型也能跨风格迁移,结果可追溯可验证

视觉语言模型在字面意义的多模态任务(如视觉数学与科学问答)中表现出色,但对讽刺、幽默、隐喻等修辞语言仍面临挑战,因其通过表意与意图间的微妙矛盾传递情感与意图。在多模态场景中,图像可能强化或反转文本含义,要求模型具备跨模态推理与主观性建模能力。本文提出一个三阶段框架,旨在构建高效多模态推理模型:(i) 解读多模态修辞语言,(ii) 提供可解释的推理轨迹,(iii) 实现多种修辞风格间的泛化。在四种修辞风格上的实验表明:(1) 引入推理轨迹显著提升多模态修辞理解性能;(2) 某一风格中学到的推理能力可有效迁移到其他风格,尤其在讽刺与幽默间迁移效果突出;(3) 联合训练多风格数据可获得超越更大规模开源与闭源模型的通用推理视觉语言模型。研究发现,具备可验证推理路径的轻量级模型,在实现稳健跨风格泛化的同时,仍能提供可审查的推理过程。代码与实现已公开于 https://github.com/scheshmi/CrossStyle-MMR。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated strong reasoning abilities in literal multimodal tasks such as visual mathematics and science question answering. However, figurative language, such as sarcasm, humor, and metaphor, remains a significant challenge, as it conveys intent and emotion through subtle incongruities between expressed and intended meanings. In multimodal settings, accompanying images can amplify or invert textual meaning, demanding models that reason across modalities and account for subjectivity. We propose a three-step framework for developing efficient multimodal reasoning models that can (i) interpret multimodal figurative language, (ii) provide transparent reasoning traces, and (iii) generalize across multiple figurative styles. Experiments across four styles show that (1) incorporating reasoning traces substantially improves multimodal figurative understanding, (2) reasoning learned in one style can transfer to others, especially between related styles like sarcasm and humor, and (3) training jointly across styles yields a generalized reasoning VLM that outperforms much larger open- and closed-source models. Our findings show that lightweight VLMs with verifiable reasoning achieve robust cross-style generalization while providing inspectable reasoning traces for multimodal tasks. The code and implementation are available at https://github.com/scheshmi/CrossStyle-MMR.

多模态推理修辞理解可解释性跨风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。