arXiv:2601.21403cs.AIcs.MA2026-01被引 5

打通结构化与视觉数据,让分析模型能读懂报表和表格。

DataCross: A Unified Benchmark and Agent Framework for Cross-Modal Heterogeneous Data Analysis

  • 设计分步协作的智能体框架,专攻跨模态数据融合。
  • 在200个真实任务中,事实准确率比GPT-4o高29.7%。
  • 适合需要处理发票、报告等非结构化数据的工业场景。

现实世界的数据科学与企业决策中,关键信息常分散在可查询的结构化数据(如SQL、CSV)和锁定在非结构化视觉文档(如扫描报告、发票图片)中的“僵尸数据”之间。现有数据分析智能体主要局限于结构化数据,无法激活并关联这些高价值的视觉信息,造成与工业需求的重大脱节。为此,我们提出DataCross,一个统一的基准与协作智能体框架,实现跨异构数据模态的洞察驱动分析。DataCrossBench包含金融、医疗等多个领域共200个端到端分析任务,通过人机协同逆向合成流程构建,确保任务具有真实复杂性、跨源依赖性和可验证真值。该基准将任务分为三个难度层级,评估智能体在视觉表格提取、跨模态对齐和多步联合推理方面的能力。我们还提出DataCrossAgent框架,借鉴人类分析师的‘分而治之’工作流,采用各司其职的子智能体,通过结构化工作流——源内深度探索、关键源识别、上下文交叉融合——进行协调。新颖的reReAct机制支持鲁棒的代码生成与调试,用于事实验证。实验表明,DataCrossAgent在事实性上比GPT-4o提升29.7%,并在高难度任务中表现更稳健,有效激活了分散的‘僵尸数据’,实现跨模态深度分析。

原文摘要 · Abstract (English)

In real-world data science and enterprise decision-making, critical information is often fragmented across directly queryable structured sources (e.g., SQL, CSV) and "zombie data" locked in unstructured visual documents (e.g., scanned reports, invoice images). Existing data analytics agents are predominantly limited to processing structured data, failing to activate and correlate this high-value visual information, thus creating a significant gap with industrial needs. To bridge this gap, we introduce DataCross, a novel benchmark and collaborative agent framework for unified, insight-driven analysis across heterogeneous data modalities. DataCrossBench comprises 200 end-to-end analysis tasks across finance, healthcare, and other domains. It is constructed via a human-in-the-loop reverse-synthesis pipeline, ensuring realistic complexity, cross-source dependency, and verifiable ground truth. The benchmark categorizes tasks into three difficulty tiers to evaluate agents' capabilities in visual table extraction, cross-modal alignment, and multi-step joint reasoning. We also propose the DataCrossAgent framework, inspired by the "divide-and-conquer" workflow of human analysts. It employs specialized sub-agents, each an expert on a specific data source, which are coordinated via a structured workflow of Intra-source Deep Exploration, Key Source Identification, and Contextual Cross-pollination. A novel reReAct mechanism enables robust code generation and debugging for factual verification. Experimental results show that DataCrossAgent achieves a 29.7% improvement in factuality over GPT-4o and exhibits superior robustness on high-difficulty tasks, effectively activating fragmented "zombie data" for insightful, cross-modal analysis.

跨模态分析智能体数据挖掘视觉表格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。