为科学可视化智能体建立评估体系,推动技术迭代与协作创新。
An Evaluation-Centric Paradigm for Scientific Visualization Agents
- 提出以评估为中心的范式,系统梳理科学可视化智能体的评测需求。
- 指出当前缺乏大规模真实场景评测基准,制约了能力比较与进步衡量。
- 倡导构建共享评估基准,助力智能体自我优化与领域发展。
多模态大语言模型的进展使自主可视化智能体能将用户意图转化为数据可视化,但在科学可视化领域,由于缺乏全面且大规模的评测基准,衡量进展与比较不同智能体仍面临挑战。本文探讨了科学可视化智能体所需的各类评估形式,分析相关难题,提供一个简单的概念验证评估示例,并讨论评测基准如何促进智能体自我改进。我们主张推动更广泛的合作,共同构建科学可视化智能体评测基准,不仅用于评估现有能力,更能驱动创新并激发未来发展方向。
原文摘要 · Abstract (English)
Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring progress and comparing different agents remains challenging, particularly in scientific visualization (SciVis), due to the absence of comprehensive, large-scale benchmarks for evaluating real-world capabilities. This position paper examines the various types of evaluation required for SciVis agents, outlines the associated challenges, provides a simple proof-of-concept evaluation example, and discusses how evaluation benchmarks can facilitate agent self-improvement. We advocate for a broader collaboration to develop a SciVis agentic evaluation benchmark that would not only assess existing capabilities but also drive innovation and stimulate future development in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。