arXiv:2606.00384cs.AIcs.CL2026-06被引 2

让AI自动设计分析工具,提升科学数据建模效率

VESTA: Visual Exploration with Statistical Tool Agents

论文配图:VESTA: Visual Exploration with Statistical Tool Agents
图 1 · 摘自论文原文
  • 用动态生成的诊断工具引导模型迭代优化
  • 在天文学真实任务中比现有方法提升显著,复杂任务增益更大
  • 适合需要深度数据分析的科研人员使用

将定量模型拟合至数据是科学工作流的核心步骤,但仍是自动化程度最低的环节。近期基于代理的系统利用语言和视觉-语言模型(VLMs)迭代提出并优化统计模型,但在更具挑战性的建模任务上表现不佳。为此,我们提出VESTA:视觉探索与统计工具代理框架,赋予VLMs一个可动态增长的探索工具集,通过数据变换、假设驱动的可视化和稳健的统计检验来指导模型优化。不同于仅依赖迭代批判的系统,VESTA在模型优化前与过程中主动探索数据,选择或创建诊断工具,这些工具积累在模型上下文中可重复使用。我们在三种工具配置下评估VESTA:无工具、静态专家编写的工具、以及动态模型生成的工具。为支持评估,我们引入DAWN(Dataset for Automated Workflows and Numerical Modeling),一个面向分布拟合与时间序列建模的基准,涵盖不同难度层级,并最终应用于真实的天文学任务,包括初始质量函数和引力波啁啾信号建模。结果表明,动态工具生成显著优于已有代理流程,在复杂且领域特定的任务上提升最明显。此外,动态生成的工具在功能覆盖度和诊断类别多样性上远超现有视觉工具生成系统,更偏好于能被VLM直接推理的可视化输出。

原文摘要 · Abstract (English)

Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems leverage language and vision-language models (VLMs) to iteratively propose and refine statistical models, but these systems struggle on more challenging modeling tasks. To address these limitations, we introduce VESTA: Visual Exploration with Statistical Tool Agents, a framework that equips VLMs with a dynamically growing exploration toolkit to guide model refinement through data transformations, hypothesis-driven visualizations, and robust statistical tests. Unlike prior systems that rely on iterative critique alone, VESTA actively explores data before and during refinement by selecting or creating diagnostic tools, which accumulate in the model's context and can be reused later. We evaluate VESTA against established baselines in three toolkit configurations: no tools, static expert-written tools, and dynamic model-written tools. To support this evaluation, we introduce DAWN (Dataset for Automated Workflows and Numerical Modeling), a benchmark targeting distribution fitting and time series modeling with varying difficulty tiers, and culminating in real-world astronomy tasks including modeling initial mass functions and gravitational-wave chirp signals. We find that VESTA's dynamic tool creation outperforms prior agentic pipelines, with the largest gains on complex and domain-specific tasks. We further show that dynamically generated tools are substantially more sophisticated than those produced by existing visual tool-creation systems, covering more diagnostic categories per function and strongly preferring visual outputs that the VLM critic can reason over directly.

科学计算智能代理数据分析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。