测试AI在物理成像任务中的表现,发现其远不如专用算法。
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

- 构建20个跨五类的计算成像任务基准,评估AI处理物理问题能力。
- 主流AI模型在无透镜成像等任务上表现显著落后于专业方法。
- 尽管输出图像看起来合理,但物理一致性差,适合研究成像可靠性者关注。
视觉语言模型和代理型AI在语义视觉任务中表现优异,但其是否能应对计算成像背后的物理规律与逆问题仍不明确。本文提出ImagingBench,一个涵盖五个类别共20个计算成像任务的基准,包括射线与波动光学、图像信号处理、逆重建、计算感知和校准。评估三种设置:专家固定引导逆重建(Expert)、规划器引导逆重建(Planner)和前向系统仿真(Forward),用于一致性验证。测试了Gemini、GPT、Qwen等主流多模态系统,并与代表性非代理型专用基线对比。结果表明,代理模型整体表现弱于专用方法,尤其在无透镜成像、事件重建、飞行时间成像和全息等领域。规划器引导仅带来微弱且不稳定的提升,优于固定提示的专家基线。尽管模型常生成视觉上合理的输出,其基于参考的保真度依然很低,揭示了语义视觉能力与物理可靠成像性能之间的巨大差距。ImagingBench为衡量这一差距并追踪代理型AI在计算成像中的进展提供了统一测试平台。
原文摘要 · Abstract (English)
Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates three complementary settings: Expert, fixed expert-guided inverse reconstruction; Planner, planner-guided inverse reconstruction; and Forward, forward-system simulation for consistency checking. We benchmark leading proprietary and open-source image-centric multimodal systems, including Gemini, GPT, and Qwen, and compare them with representative task-specific non-agentic baselines. Across tasks, agentic models remain consistently weaker than specialized methods, especially on computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Planner guidance provides only modest and inconsistent gains over the fixed-prompt Expert baseline. Although the models often generate visually plausible outputs, their reference-based fidelity remains poor, revealing a substantial gap between semantic visual competence and physically grounded imaging performance. ImagingBench provides a unified testbed for measuring this gap and tracking progress in agentic AI for computational imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。