评测智能体在真实金融数据中自主探索分析的能力
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis

- 构建真实金融场景下的多领域噪声数据环境,模拟无指引探索
- 包含206万条记录和492个跨领域任务,支持过程诊断与失败分析
- 揭示当前大模型探索越多越难出正确结论,暴露可靠性短板
自主数据分析智能体被期待在极少人为干预下完成探索性分析。然而,现有基准大多在预设引导环境下评估,提供选定数据源、明确数据模式或清洗后的数据,低估了真实探索的难度。为评估这一现实任务,我们提出DataClawBench,一个基于金融智库咨询场景的基准,要求智能体独立探索陌生、嘈杂、跨领域的数据并生成可验证结论。该基准提供约206万条记录的统一真实数据环境,涵盖企业、行业与政策领域,保留原始数据噪声。在此基础上定义492个多步骤跨领域任务,每项任务配备中间里程碑标注,可诊断探索与推理失败,不仅关注最终结果准确性。对八种先进大模型在OpenClaw智能体框架下的系统评估显示,探索行为并不必然带来任务进展或正确答案,揭示探索性数据分析严重削弱了智能体的可靠性。
原文摘要 · Abstract (English)
Autonomous data analysis agents are increasingly expected to conduct exploratory analysis with limited human guidance about data. However, existing benchmarks typically evaluate such agents in prior-guided settings, providing selected data sources, explicit data schemas, or cleaned data, thereby understating the exploratory burden. To evaluate this realistic exploratory data analysis task, we introduce DataClawBench, a benchmark built from financial think-tank consulting scenarios where agents must independently explore unfamiliar, noisy, cross-domain data and produce verifiable conclusions. DataClawBench provides a unified real-world data environment with approximately 2.06 million records across enterprise, industry, and policy domains, with native data noise preserved. On top of this data environment, it defines 492 multi-step cross-domain tasks, each annotated with intermediate milestones that diagnose exploration and reasoning failures beyond outcome accuracy. A systematic evaluation of eight advanced LLMs under the OpenClaw agent reveals that exploratory data analysis breaks agent reliability: more exploration does not reliably translate into task-relevant progress or correct final answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。