评测大模型在科学计算成像任务中的代码生成能力,揭示其在物理建模与流程整合上的系统性短板。
Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging

- 构建57个跨六领域的标准化成像任务基准,含预处理、正向建模、反演求解与可视化四阶段
- 7个前沿大模型在端到端重建中表现不佳,尤其在算法选择与物理惯例处理上问题突出
- 适合研究计算成像、提示工程与领域专用智能体的开发者参考
计算成像通过从间接、噪声测量中恢复隐藏信号,支撑多个科学领域的定量发现,但构建正确重建流程需深厚领域知识,即使对领域科学家也极为耗时。本文提出Imaging-101,一个包含57个专家验证的计算成像任务基准,覆盖六个科学领域,每个任务基于同行评审论文并统一为标准化四阶段流程(预处理、正向物理建模、逆向求解器、可视化)。三个评估赛道(规划、函数级单元测试、端到端重建)分别考察代理在全流程中的不同能力。对七个前沿大语言模型的评估揭示了编码代理应用于计算成像时系统性挑战,超越通用编码基准暴露的问题,涵盖算法选择、物理惯例处理与流程集成等维度。这些发现凸显了具体能力缺口,并指向技能增强、领域专化的智能体是实现可靠计算成像辅助的可行路径。
原文摘要 · Abstract (English)
Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious even for domain scientists. We introduce Imaging-101, a benchmark of 57 expert-verified computational imaging tasks spanning six scientific domains, each grounded in a peer-reviewed paper and canonicalized into a standardized four-stage pipeline (preprocessing, forward physics modeling, inverse solver, and visualization) Three evaluation tracks (planning, function-level unit tests, and end-to-end reconstruction) probe distinct agent capabilities across the full pipeline. Evaluating seven frontier LLMs uncovers systematic challenges in applying coding agents to computational imaging that go beyond those exposed by general coding benchmarks, spanning algorithm selection, physical convention handling, and pipeline integration. These findings highlight concrete capability gaps and point toward skill-augmented, domain-specialized agents as a practical path to reliable computational imaging assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。