arXiv:2411.07664cs.CV2024-11被引 8

对比图文生成模型与大语言模型的空间关系理解能力,发现后者更准确。

Evaluating the Generation of Spatial Relations in Text and Image Generative Models

  • 将大模型文本输出转为图像,实现对两类模型的可视化评估
  • 8个模型在10个介词上的空间关系生成准确率普遍不高,尤其图文模型表现差
  • 大模型虽以文本训练为主,但空间理解优于图文生成模型,适合研究空间认知

空间关系理解是人类与AI的重要认知能力。当前研究多聚焦于文本到图像(T2I)模型的评测,本文提出更全面的评估框架,涵盖T2I模型和大语言模型(LLMs)。由于空间关系具有视觉空间特性,我们设计方法将LLM输出转换为图像,从而实现对两类模型的可视化评估。我们在10个常见介词上测试了8个主流生成模型(3个T2I模型、5个LLMs)的空间关系生成能力,并评估自动评测方法的可行性。结果显示,尽管T2I模型具备强大的通用图像生成能力,其空间关系生成表现却较差;更意外的是,LLMs在空间关系生成上显著优于T2I模型,尽管它们主要基于文本数据训练。我们分析了模型失败原因,指出可填补的差距,以提升生成结果的空间一致性。

原文摘要 · Abstract (English)

Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) models, we propose a more comprehensive evaluation that includes \textit{both} T2I and Large Language Models (LLMs). As spatial relations are naturally understood in a visuo-spatial manner, we develop an approach to convert LLM outputs into an image, thereby allowing us to evaluate both T2I models and LLMs \textit{visually}. We examined the spatial relation understanding of 8 prominent generative models (3 T2I models and 5 LLMs) on a set of 10 common prepositions, as well as assess the feasibility of automatic evaluation methods. Surprisingly, we found that T2I models only achieve subpar performance despite their impressive general image-generation abilities. Even more surprisingly, our results show that LLMs are significantly more accurate than T2I models in generating spatial relations, despite being primarily trained on textual data. We examined reasons for model failures and highlight gaps that can be filled to enable more spatially faithful generations.

空间关系图文生成大语言模型评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。