通过隐含世界重建评估企业数据智能体的分析理解能力。
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
- 以恢复数据背后的分析结构为核心,而非仅验证流程执行。
- 在电商场景中,最强代码代理仅正确恢复26%的分析要素。
- 可追踪早期分析错误如何导致后续结论系统性偏差。
我们提出AvalancheBench,一种基于隐含世界重建的企业数据智能体评估基准。该基准在三方面超越现有方法:其一,评估分析理解而非流程完成度——系统得分取决于是否恢复解释数据的细分群体、驱动因素、时间事件及关系;其二,通过已知隐含世界生成观测数据,实现对部分有效恢复的局部评分;其三,揭示早期分析错误(如遗漏群体、合并事件、错误归因)如何导致后续结论系统性偏差。AvalancheBench为诊断智能体是否准确还原企业数据背后分析结构提供可控环境。在首个电商应用场景中,顶尖编码智能体仅恢复26%的基准要素,失败主要集中在通用客户分群和时间事件合并上。
原文摘要 · Abstract (English)
We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through \emph{latent world recovery}. AvalancheBench improves on existing benchmarks in three ways. First, it evaluates analytical understanding rather than pipeline completion: systems are scored on whether they recover the segments, drivers, temporal events, and relationships that explain the data, not merely on whether they execute a workflow or produce a plausible report. Second, it provides ground truth for goal-driven analytics by generating observations from a known latent world, enabling partial credit for incomplete but valid recoveries. Third, it exposes how early analytical mistakes propagate into later conclusions: missed segments, merged events, or wrong attributions can lead to systematically wrong recommendations. In this sense, AvalancheBench complements real-data benchmarks by providing a controlled setting for diagnosing whether agents recover the analytical structure behind enterprise data. On a first e-commerce use case, the strongest configuration of a leading coding agent recovers only 26\% of the rubric, with failures concentrated in generic customer segmentations and merged temporal events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。