用照片测试大模型的视觉推理能力,看它能否猜出相机设置。
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
- 基于照片中光影模糊等物理特征,推断具体相机参数。
- 现有模型在不同任务间表现不一,无统一优势。
- 适合研究视觉理解与摄影智能助手的开发者参考。
大型语言模型(LLMs)和多模态大语言模型(MLLMs)显著推动了人工智能发展。然而,涉及视觉与文本双重输入的视觉推理仍待深入探索。近期如 OpenAI o1 与 Gemini 2.0 Flash Thinking 等推理模型引入图像输入,开启了这一能力。本文聚焦摄影相关任务,因为照片是物理世界的视觉快照,其成像受光照、模糊程度等底层物理规律与相机参数共同影响。从照片的视觉信息准确推断出具体的数值化相机设置,要求 MLLMs 具备对底层物理机制的深刻理解,属于高阶且实用的智能能力,适用于摄影助手等应用场景。我们旨在评估 MLLMs 在区分与数值相机设置相关的视觉差异方面的表现,并扩展此前用于视觉-语言模型(VLMs)的方法。初步结果表明,视觉推理在摄影任务中至关重要,且当前没有单一模型在所有任务中持续领先,凸显了提升 MLLMs 视觉推理能力的挑战与机遇。
原文摘要 · Abstract (English)
Large language models (LLMs) and multimodal large language models (MLLMs) have significantly advanced artificial intelligence. However, visual reasoning, reasoning involving both visual and textual inputs, remains underexplored. Recent advancements, including the reasoning models like OpenAI o1 and Gemini 2.0 Flash Thinking, which incorporate image inputs, have opened this capability. In this ongoing work, we focus specifically on photography-related tasks because a photo is a visual snapshot of the physical world where the underlying physics (i.e., illumination, blur extent, etc.) interplay with the camera parameters. Successfully reasoning from the visual information of a photo to identify these numerical camera settings requires the MLLMs to have a deeper understanding of the underlying physics for precise visual comprehension, representing a challenging and intelligent capability essential for practical applications like photography assistant agents. We aim to evaluate MLLMs on their ability to distinguish visual differences related to numerical camera settings, extending a methodology previously proposed for vision-language models (VLMs). Our preliminary results demonstrate the importance of visual reasoning in photography-related tasks. Moreover, these results show that no single MLLM consistently dominates across all evaluation tasks, demonstrating ongoing challenges and opportunities in developing MLLMs with better visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。