arXiv:2602.10575cs.CVcs.AI2026-02被引 2

用强化学习让AI读懂图像隐喻,性能提升82.6%。

MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning

  • 构建端到端视觉强化学习框架,实现图像隐喻理解
  • 在多个任务上平均提升82.6%,超越20+主流模型
  • 适合研究视觉推理、文化语义与模型泛化能力的学者

图像隐喻理解仍是当前AI系统的关键挑战。尽管多模态大语言模型在基础视觉问答中表现优异,却难以把握视觉内容中蕴含的文化、情感与上下文深层含义。这一难题源于任务所需的多跳推理、文化背景与心智理论能力,现有模型普遍缺失。为此,我们提出首个端到端视觉强化学习框架MetaphorStar,包含细粒度数据集TFQ-Data、视觉强化学习方法TFQ-GRPO及结构化基准TFQ-Bench。基于TFQ-Data训练的MetaphorStar全开源系列,在图像隐喻基准上平均性能提升82.6%。相比20余种主流多模态大模型,MetaphorStar-32B在选择题与开放题任务达到最新水平,显著优于闭源模型Gemini-3.0-pro在判断题上的表现。实验还表明,学习图像隐喻能有效提升模型的通用理解能力,尤其增强复杂视觉推理能力。我们进一步系统分析了参数量、数据量、模型架构与训练策略的影响,验证方法的广泛适用性。所有模型权重、数据集与代码已开源:https://metaphorstar.github.io。

原文摘要 · Abstract (English)

Metaphorical comprehension in images remains a critical challenge for Nowadays AI systems. While Multimodal Large Language Models (MLLMs) excel at basic Visual Question Answering (VQA), they consistently struggle to grasp the nuanced cultural, emotional, and contextual implications embedded in visual content. This difficulty stems from the task's demand for sophisticated multi-hop reasoning, cultural context, and Theory of Mind (ToM) capabilities, which current models lack. To fill this gap, we propose MetaphorStar, the first end-to-end visual reinforcement learning (RL) framework for image implication tasks. Our framework includes three core components: the fine-grained dataset TFQ-Data, the visual RL method TFQ-GRPO, and the well-structured benchmark TFQ-Bench. Our fully open-source MetaphorStar family, trained using TFQ-GRPO on TFQ-Data, significantly improves performance by an average of 82.6% on the image implication benchmarks. Compared with 20+ mainstream MLLMs, MetaphorStar-32B achieves state-of-the-art (SOTA) on Multiple-Choice Question and Open-Style Question, significantly outperforms the top closed-source model Gemini-3.0-pro on True-False Question. Crucially, our experiments reveal that learning image implication tasks improves the general understanding ability, especially the complex visual reasoning ability. We further provide a systematic analysis of model parameter scaling, training data scaling, and the impact of different model architectures and training strategies, demonstrating the broad applicability of our method. We open-sourced all model weights, datasets, and method code at https://metaphorstar.github.io.

图像理解隐喻推理强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。