测试AI在复杂空间生物数据中推导科学结论的能力
Verifiable Benchmarking of Long-Horizon Spatial Biology
- 构建跨多模态空间数据的长时程生物推理基准
- 三组模型组合在72次运行中均达成11.1%正确率
- 适合评估AI在真实生物研究中的推理能力
AI代理在生物数据分析中日益重要,但现有基准多聚焦于广泛生物学知识、可执行工作流或局部分析步骤,而非对空间测量数据的端到端科学推理。我们提出SpatialBench-Long,一个针对长时程空间生物学的基准,要求代理从原始或近原始数据及校准实验背景中恢复生物结论,无需预设方法。该基准涵盖胰腺导管腺癌(PDAC)、工程化胶质母细胞瘤类器官、体内肿瘤、Cas9谱系追踪肺腺癌及小鼠视神经衰老/干预系统,覆盖CosMx、Visium、Xenium、MERFISH、scRNA-seq、Slide-seq、Slide-tags、组织病理学与谱系记录数据。候选结论经重复验证、独立科学家评审与轨迹检查加固。最终答案基于受控词汇表和符号进行确定性评分,配套评分标准捕捉关键分析瓶颈进展。在SpatialBench-Long基准中,三组模型-工具组合在72次运行中均取得8分(11.1%),分别为Gemini 3.5 Flash / Pi终端编码工具、GPT-5.5 / Pi、GPT-5.5 / OpenAI Codex。该基准测试代理能否超越流程化分析,从复杂空间测量中得出准确科学结论。
原文摘要 · Abstract (English)
AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps rather than end-to-end scientific reasoning over spatial measurements. We introduce SpatialBench-Long, a benchmark for long-horizon spatial biology in which agents must recover biological claims from raw or near-raw data and calibrated experimental context without prescribed methods. SpatialBench-Long contains 24 evaluations across primary pancreatic ductal adenocarcinoma (PDAC), engineered glioblastoma organoids and in vivo tumors, Cas9 lineage-traced lung adenocarcinoma, and mouse optic nerve aging/intervention systems, spanning CosMx, Visium, Xenium, multiplexed error-robust fluorescence in situ hybridization (MERFISH), single-cell RNA sequencing (scRNA-seq), Slide-seq, Slide-tags, histology, and lineage-recording data. Candidate claims are hardened through reproduction, independent scientist review, and trajectory inspection. Final answers are graded deterministically over controlled vocabularies and symbols with companion rubrics capturing progress through key analysis chokepoints. Across the SpatialBench-Long benchmark, three model-harness pairs tie at 8/72 runs (11.1\%): Gemini 3.5 Flash / Pi terminal coding harness, GPT-5.5 / Pi, and GPT-5.5 / OpenAI Codex. SpatialBench-Long tests whether agents can move beyond executing procedural analysis to deriving accurate scientific conclusions from complex spatial measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。