用大模型代理实现科研工作流的实时数据溯源分析
LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology
- 通过自然语言转结构化查询,让大模型动态解析工作流元数据
- 在真实化学工作流上验证,多模型响应准确率显著高于传统方法
- 适合需要可复现、可解释科研分析的研究者和工具开发者
现代科学发现越来越多地依赖于跨边缘、云与高性能计算(HPC)环境的数据处理工作流。对这些数据进行深入分析对于假设验证、异常检测、可复现性及重要发现至关重要。尽管工作流溯源技术支持此类分析,但大规模下溯源数据变得复杂难解。现有系统依赖定制脚本、结构化查询或静态仪表板,限制了数据交互性。本文提出一种评估方法、参考架构及开源实现,利用交互式大语言模型(LLM)代理进行运行时数据分析。该方法采用轻量级元数据驱动设计,将自然语言转化为结构化溯源查询。在LLaMA、GPT、Gemini和Claude上进行评估,涵盖多种查询类型及一个真实化学工作流,结果表明模块化设计、提示调优与检索增强生成(RAG)使LLM代理能生成超越原始溯源信息的准确且有洞察力的回答。
原文摘要 · Abstract (English)
Modern scientific discovery increasingly relies on workflows that process data across the Edge, Cloud, and High Performance Computing (HPC) continuum. Comprehensive and in-depth analyses of these data are critical for hypothesis validation, anomaly detection, reproducibility, and impactful findings. Although workflow provenance techniques support such analyses, at large scale, the provenance data become complex and difficult to analyze. Existing systems depend on custom scripts, structured queries, or static dashboards, limiting data interaction. In this work, we introduce an evaluation methodology, reference architecture, and open-source implementation that leverages interactive Large Language Model (LLM) agents for runtime data analysis. Our approach uses a lightweight, metadata-driven design that translates natural language into structured provenance queries. Evaluations across LLaMA, GPT, Gemini, and Claude, covering diverse query classes and a real-world chemistry workflow, show that modular design, prompt tuning, and Retrieval-Augmented Generation (RAG) enable accurate and insightful LLM agent responses beyond recorded provenance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。