arXiv:2505.17019cs.CVcs.AI2025-05被引 6

让AI像人一样理解图像隐喻,突破视觉语义认知瓶颈。

Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework

  • 分三阶段模拟人类认知:感知、搜索、推理,补全图像上下文
  • 在中英文隐喻理解任务上均达顶尖水平,中文提升显著
  • 轻量模型媲美顶级大模型,适用于通用视觉问答与推理

图像隐喻理解仍是AI系统的核心挑战,现有模型难以捕捉视觉内容中的文化、情感与上下文深层含义。尽管多模态大语言模型(MLLMs)在通用视觉问答(VQA)任务中表现优异,但在图像隐喻任务上存在根本性局限:上下文缺失导致视觉元素间抽象关系不明确。受人类认知过程启发,我们提出Let Androids Dream(LAD)框架,通过三阶段机制解决上下文缺失问题:(1)感知阶段将视觉信息转化为多层次文本表示;(2)搜索阶段迭代检索并融合跨领域知识以消除歧义;(3)推理阶段通过显式推理生成上下文一致的图像隐喻理解。采用轻量级GPT-4o-mini模型,LAD在英文隐喻基准上超越15+ MLLMs,中文基准性能大幅提升,多项指标接近Gemini-3.0-pro,开放问答(OSQ)任务上优于GPT-4o 36.7%。泛化实验表明,该框架亦能有效提升通用VQA与视觉推理能力。本工作为人工智能更有效地解析图像隐喻提供了新视角,推动视觉语言推理与人机交互发展。项目开源地址:https://github.com/MING-ZCH/Let-Androids-Dream-of-Electric-Sheep。

原文摘要 · Abstract (English)

Metaphorical comprehension in images remains a critical challenge for AI systems, as existing models struggle to grasp the nuanced cultural, emotional, and contextual implications embedded in visual content. While multimodal large language models (MLLMs) excel in general Visual Question Answer (VQA) tasks, they struggle with a fundamental limitation on image implication tasks: contextual gaps that obscure the relationships between different visual elements and their abstract meanings. Inspired by the human cognitive process, we propose Let Androids Dream (LAD), a novel framework for image implication understanding and reasoning. LAD addresses contextual missing through the three-stage framework: (1) Perception: converting visual information into rich and multi-level textual representations, (2) Search: iteratively searching and integrating cross-domain knowledge to resolve ambiguity, and (3) Reasoning: generating context-alignment image implication via explicit reasoning. Our framework with the lightweight GPT-4o-mini model achieves SOTA performance compared to 15+ MLLMs on English image implication benchmark and a huge improvement on Chinese benchmark, performing comparable with the Gemini-3.0-pro model on Multiple-Choice Question (MCQ) and outperforms the GPT-4o model 36.7% on Open-Style Question (OSQ). Generalization experiments also show that our framework can effectively benefit general VQA and visual reasoning tasks. Additionally, our work provides new insights into how AI can more effectively interpret image implications, advancing the field of vision-language reasoning and human-AI interaction. Our project is publicly available at https://github.com/MING-ZCH/Let-Androids-Dream-of-Electric-Sheep.

图像理解隐喻推理多模态视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。