arXiv:2603.12483cs.AIcs.LG2026-03被引 1

为时序数据分析代理定制高效评估,解决真实场景下的表达力短板

Generating Expressive and Customizable Evals for Timeseries Data Analysis Agents with AgentFuel

  • 基于领域特性的数据与查询类型构建可定制评估框架
  • 在6个主流代理上测试,发现其在状态相关和事件类查询中表现差
  • 适合需要真实场景验证的工业级数据代理开发者

在物联网、可观测性、通信与网络安全等多个领域,对话式数据分析代理正被广泛采用,使用户能够“与数据对话”获取洞察。这些代理处理时序数据模型,如传感器测量值或产品分析中的用户点击事件。我们对6个主流数据代理(含开源与专有)在特定领域数据和查询类型上进行评估,发现它们在状态依赖和事件相关的查询上表现不佳。现有评估存在两大关键缺陷:缺乏领域定制的数据集和领域特定的查询类型。为此,我们提出AgentFuel,帮助领域专家快速构建定制化评估,实现端到端功能测试。实验表明,AgentFuel暴露了现有数据代理框架的关键改进方向。此外,我们还获得初步证据显示使用AgentFuel可提升代理性能(如GEPA)。AgentFuel基准数据集已公开于https://huggingface.co/datasets/RockfishData/TimeSeriesAgentEvals。

原文摘要 · Abstract (English)

Across many domains (e.g., IoT, observability, telecommunications, cybersecurity), there is an emerging adoption of conversational data analysis agents that enable users to "talk to your data" to extract insights. Such data analysis agents operate on timeseries data models; e.g., measurements from sensors or events monitoring user clicks and actions in product analytics. We evaluate 6 popular data analysis agents (both open-source and proprietary) on domain-specific data and query types, and find that they fail on stateful and incident-specific queries. We observe two key expressivity gaps in existing evals: domain-customized datasets and domain-specific query types. To enable practitioners in such domains to generate customized and expressive evals for such timeseries data agents, we present AgentFuel. AgentFuel helps domain experts quickly create customized evals to perform end-to-end functional tests. We show that AgentFuel's benchmarks expose key directions for improvement in existing data agent frameworks. We also present anecdotal evidence that using AgentFuel can improve agent performance (e.g., with GEPA). AgentFuel benchmarks are available at https://huggingface.co/datasets/RockfishData/TimeSeriesAgentEvals.

数据代理时序分析评估框架领域定制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。