首个评估多模态模型光学生成与理解能力的基准测试
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
- 构建几何光学场景提示,系统评估模型生成真实图像能力
- 生成模型最高仅完成部分任务,理解模型准确率不足37.4%
- 适合关注视觉物理真实性与多模态模型局限的研究者
多模态大模型在视觉理解与生成方面进展迅速,但对其在几何光学等细粒度物理原理上的能力评估仍不充分。为此,我们提出GOBench,首个系统评估多模态大模型在两大任务上的基准:1)生成符合光学规律的图像;2)理解底层光学现象。我们构建了高质量的几何光学场景提示,并利用多模态大模型生成GOBench-Gen-1k数据集。通过主观实验评估生成图像的光学真实性、美学质量与指令一致性,揭示模型生成中违反光学原理的问题。在理解任务中,对十一款主流多模态大模型使用精心设计的评测指令进行测试。结果表明,当前模型在光学生成与理解上均面临显著挑战:表现最佳的生成模型GPT-4o-Image未能完全完成所有任务,而表现最佳的理解模型Gemini-2.5Pro仅达到37.35%的准确率。数据集与代码已开源。
原文摘要 · Abstract (English)
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curates high-quality prompts of geometric optical scenarios and use MLLMs to construct GOBench-Gen-1k dataset.We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35\% accuracy in optical understanding. Database and codes are publicly available at https://github.com/aiben-ch/GOBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。