arXiv:2603.29139cs.AIcs.GR2026-03被引 9

首个面向科学可视化智能体的综合性评估基准,支持多步分析任务测试。

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

  • 构建四维分类体系,覆盖108个专家设计的真实科研场景。
  • 融合多模态评估机制,准确率超90%且与人工评判高度一致。
  • 适合研究科学计算、AI代理和数据可视化的开发者与学者使用。

大型语言模型的进展使得智能体能够将自然语言意图转化为可执行的科学可视化任务。然而,社区仍缺乏一个系统化且可复现的基准来评估这些新兴的科学可视化智能体在真实、多步骤分析环境中的表现。本文提出SciVisAgentBench,一个全面且可扩展的基准,用于评估科学数据分析与可视化智能体。该基准基于涵盖应用领域、数据类型、复杂度水平和可视化操作四个维度的结构化分类体系,包含108个由专家精心设计的案例,覆盖多样化的科学可视化场景。为实现可靠评估,我们引入一种以结果为中心的多模态评估流程,结合基于LLM的判别方法与确定性评估工具,包括图像度量、代码检查器、规则验证器及针对特定案例的评估器。我们还通过12位科学可视化专家开展有效性研究,检验人类与LLM判别者的一致性。利用此框架,我们对代表性科学可视化智能体及通用编程智能体进行评估,建立初始基线并揭示能力差距。SciVisAgentBench设计为持续演进的活基准,支持系统性对比、故障模式诊断与领域进步。基准代码与数据已公开:https://scivisagentbench.github.io/。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We also conduct a validity study with 12 SciVis experts to examine the agreement between human and LLM judges. Using this framework, we evaluate representative SciVis agents and general-purpose coding agents to establish initial baselines and reveal capability gaps. SciVisAgentBench is designed as a living benchmark to support systematic comparison, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is available at https://scivisagentbench.github.io/.

科学可视化智能体评估基准测试LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。