测试AI生成图像时对3D空间关系的理解能力,发现模型常出错。
GenSpace: Benchmarking Spatially-Aware Image Generation
- 用多模型重建3D场景几何,评估图像空间准确性
- 现有模型在物体位置、关系和尺寸上错误频发
- 适合研究视觉生成与空间认知的学者参考
人类能直观地在三维空间中构图摄影,但先进的AI图像生成模型在根据文本或图像提示生成图像时,是否具备类似的3D空间感知能力?我们提出GenSpace,一个全新的基准测试与评估流程,全面评估当前图像生成模型的空间意识。传统基于通用视觉语言模型(VLM)的评估常无法捕捉细微的空间错误。为此,我们设计了一套专用评估流程与指标,利用多个视觉基础模型重建3D场景几何,提供更准确且符合人类直觉的空间忠实度度量。结果表明,尽管这些模型生成的图像视觉上令人满意并能遵循一般指令,但在物体放置、空间关系和度量等具体3D细节上表现不佳。我们总结出现有最先进图像生成模型在空间感知上的三大核心局限:1)物体视角理解不足;2)自我中心与环境坐标转换困难;3)度量一致性难以保持,为提升图像生成的空间智能指明了方向。
原文摘要 · Abstract (English)
Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace, a novel benchmark and evaluation pipeline to comprehensively assess the spatial awareness of current image generation models. Furthermore, standard evaluations using general Vision-Language Models (VLMs) frequently fail to capture the detailed spatial errors. To handle this challenge, we propose a specialized evaluation pipeline and metric, which reconstructs 3D scene geometry using multiple visual foundation models and provides a more accurate and human-aligned metric of spatial faithfulness. Our findings show that while AI models create visually appealing images and can follow general instructions, they struggle with specific 3D details like object placement, relationships, and measurements. We summarize three core limitations in the spatial perception of current state-of-the-art image generation models: 1) Object Perspective Understanding, 2) Egocentric-Allocentric Transformation and 3) Metric Measurement Adherence, highlighting possible directions for improving spatial intelligence in image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。