评测AI处理复杂文档分析任务的综合基准,真实场景难度高。
AIDABench: AI Data Analytics Benchmark
- 构建涵盖600+任务的端到端评估框架,覆盖问答、可视化与文件生成。
- 顶尖模型在复杂任务上仅达59.43%准确率,体现当前AI能力瓶颈。
- 适合企业选型、模型优化及研究挑战分析,推动真实场景应用发展。
随着AI驱动的文档理解与处理工具在实际应用中日益普及,建立严格的评估标准变得尤为紧迫。现有基准多聚焦单一能力或简化场景,难以反映真实世界中的全流程任务效果。为此,我们提出AIDABench,一个面向复杂数据分析任务的综合性评估基准,支持端到端评测。该基准包含600多个跨行业的真实任务,覆盖问答、数据可视化和文件生成三大核心能力维度,数据类型包括电子表格、数据库、财务报告与运营记录等。任务设计极具挑战性,即使人类专家借助AI工具,每题平均仍需1-2小时完成,凸显其现实复杂度。我们在AIDABench上评估了11个先进模型,涵盖专有(如Claude Sonnet 4.5、Gemini 3 Pro Preview)与开源(如Qwen3-Max-2026-01-23-Thinking)模型。结果表明,当前系统在复杂真实任务上仍面临重大挑战,最佳模型仅实现59.43% pass-at-1。我们进一步分析各维度的失败模式,识别出未来研究的关键挑战。AIDABench为企事业单位采购、工具选型与模型优化提供可靠参考,项目已公开于https://github.com/MichaelYang-lyx/AIDABench。
原文摘要 · Abstract (English)
As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on isolated capabilities or simplified scenarios, failing to capture the end-to-end task effectiveness required in practical settings. To address this gap, we introduce AIDABench, a comprehensive benchmark for evaluating AI systems on complex data analytics tasks in an end-to-end manner. AIDABench encompasses 600+ diverse document analysis tasks across three core capability dimensions: question answering, data visualization, and file generation. These tasks are grounded in realistic scenarios involving heterogeneous data types, including spreadsheets, databases, financial reports, and operational records, and reflect analytical demands across diverse industries and job functions. Notably, the tasks in AIDABench are sufficiently challenging that even human experts require 1-2 hours per question when assisted by AI tools, underscoring the benchmark's difficulty and real-world complexity. We evaluate 11 state-of-the-art models on AIDABench, spanning both proprietary (e.g., Claude Sonnet 4.5, Gemini 3 Pro Preview) and open-source (e.g., Qwen3-Max-2026-01-23-Thinking) families. Our results reveal that complex, real-world data analytics tasks remain a significant challenge for current AI systems, with the best-performing model achieving only 59.43% pass-at-1. We provide a detailed analysis of failure modes across each capability dimension and identify key challenges for future research. AIDABench offers a principled reference for enterprise procurement, tool selection, and model optimization, and is publicly available at https://github.com/MichaelYang-lyx/AIDABench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。