arXiv:2512.21907cs.AI2025-12被引 10

测试AI代理分析真实空间生物学数据的能力

SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

  • 构建146个真实分析任务的基准测试集
  • 主流模型准确率仅20%-38%,表现受平台和任务影响大
  • 强调提示设计与执行环境对性能的关键作用

空间转录组学数据正快速增加规模与复杂性,计算分析已成为生物发现的主要瓶颈。尽管前沿AI代理在软件工程和通用数据分析方面取得显著进步,但其能否从混乱的真实空间数据中提取生物洞见仍不明确。我们提出SpatialBench,一个涵盖五种空间技术与七类任务的基准,包含146个可验证的问题。每个问题提供分析前的数据快照及确定性评分器,用于评估关键生物结果的恢复情况。基准测试显示,基础模型准确率在各模型族间仅为20%-38%,且存在显著的模型-任务与模型-平台交互效应。提示设计对性能有明显影响,表明工具、提示、控制流与执行环境应作为核心要素进行评估与优化。SpatialBench不仅可用作测量工具,也可作为诊断视角,推动能忠实、透明、可复现地交互真实空间数据的智能代理发展。

原文摘要 · Abstract (English)

Spatial transcriptomics assays are rapidly increasing in scale and complexity, making computational analysis a major bottleneck in biological discovery. Although frontier AI agents have improved dramatically at software engineering and general data analysis, it remains unclear whether they can extract biological insight from messy, real-world spatial datasets. We introduce SpatialBench, a benchmark of 146 verifiable problems derived from practical spatial analysis workflows spanning five spatial technologies and seven task categories. Each problem provides a snapshot of experimental data immediately prior to an analysis step and a deterministic grader that evaluates recovery of a key biological result. Benchmark data on frontier models shows that base model accuracy remains low (20-38% across model families), with strong model-task and model-platform interactions. Harness design has a large empirical effect on performance, indicating that tools, prompts, control flow, and execution environment should be evaluated and improved as first-class objects. SpatialBench serves both as a measurement tool and a diagnostic lens for developing agents that can interact with real spatial datasets faithfully, transparently, and reproducibly.

空间转录组AI代理生物信息学基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。