arXiv:2608.03451cs.AI2026-08被引 1

构建跨模态数据空间基准,评估智能体在复杂环境中的可验证分析能力。

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

论文配图:DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
图 1 · 摘自论文原文
  • 通过多模态数据融合与任务驱动生成,实现自然语言到结构化表格的端到端输出。
  • 6大主流模型最佳准确率达66.34%,但不同框架导致性能差距达15.36个百分点。
  • 聚焦真实工作场景,适合研究数据智能体可靠性和多源证据整合的开发者。

数据智能体可在组织工作空间中实现自然语言分析,其中相关证据可能分散在数据库、结构化文件、长文档和多媒体中。现有基准大多孤立地评估结构化查询、检索或开放性分析,未能统一异构证据发现、完整表格输出及确定性评估。我们提出DataSpace,一个数据智能体从局部异构工作空间生成可验证表格结果的基准。该基准包含410个跨语言任务和7,439个数据资源,总大小15.01 GB,涵盖CSV、JSON、SQLite、Markdown、PDF和视频格式。DataSpace还作为KDD Cup 2026数据智能体复杂数据分析竞赛的官方评估基准。每个智能体仅接收问题和工作空间,返回完整请求的表格结果。我们使用DataSpace-Builder框架构建,该框架包括跨语言转换、约束感知关系采样、模态路由与资源渲染,以及11位领域专家的人工审核与任务修复。采用确定性评估器进行不区分表头的列对齐、类型与精度敏感的归一化及行序感知比较。在六种前沿多模态模型与五种常用代理框架下,最高准确率为66.34%,而固定主干模型时,框架选择导致15.36个百分点的性能差异。多模态证据融合与连接操作在所有六种主干模型上均显著降低准确率。这些结果表明DataSpace尚未饱和,揭示了提升数据智能体可靠性的重要挑战。

原文摘要 · Abstract (English)

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.

数据智能体多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。