构建首个需证据支撑的多模态隐喻理解基准,解决模型依赖文字表层线索的问题。
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

- 提出统一标注框架,要求模型识别视觉与文本证据并建立跨模态映射
- 实测现有模型严重忽略视觉信息,准确率低于人类水平30%以上
- 新方法用分阶段强化学习提升推理对齐度,8B小模型超越多个大模型
隐喻通过跨域映射帮助理解抽象概念并传递情感态度。在多模态场景中,视觉与文本信息共同构建目标-源映射,需要兼具概念理解与跨模态推理能力。然而,现有基准主要评估孤立子任务,缺乏证据支撑的解释,难以判断模型是否基于视觉与文本线索建立有效映射。为此,我们提出M$^3$R-Bench,一个统一且基于证据的基准,包含1,000个图像-文本实例及人工验证标注。该基准基于概念隐喻理论与非字面语言理解理论,提供隐喻存在性、目标-源映射、情感及分阶段解释(证据识别→映射建立→情感推断)的联合标注。在该基准上的评估显示,现有模型常忽略视觉证据,依赖表面文本线索,产生不准确的目标-源映射,暴露跨模态证据-映射不一致问题。为缓解此问题,我们提出M$^3$R-Reasoner,结合基于课程的推理监督与任务感知强化学习,使模型推理与隐喻解释对齐。实验表明,仅使用8B参数骨干网络,M$^3$R-Reasoner在四项统一任务指标上优于更大规模专有多模态大模型,并在视觉证据与情感论证得分上分别超过GPT-5.5达28.45和30.11点,同时在平均评分上领先Claude-Sonnet-4.6达8.00点。数据集与代码已公开于https://github.com/hongshi4/M3R-Bench。
原文摘要 · Abstract (English)
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。