用大模型自动分析结构化数据库,生成有逻辑的分析故事。
DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling
- 基于探索性数据分析与叙事思维,构建课题驱动的研究流程
- 在 InsightBench 上提升洞察召回率19.4%,摘要得分提升7.2%
- 适用于需深度挖掘数据库的科研、商业分析场景
大型语言模型(LLM)代理在多步信息发现、整合与分析中展现出强大潜力。然而,现有方法主要聚焦非结构化网络数据,对大规模结构化数据库的深度研究仍缺乏探索。与网络研究不同,数据驱动研究不仅需要检索与摘要,更需迭代生成假设、进行结构化模式的定量推理,并最终形成连贯的分析叙事。本文提出 DataSTORM,一个基于 LLM 的智能体系统,可自主跨结构化数据库与互联网源开展研究。该系统融合探索性数据分析与数据叙事原则,将深度研究重构为以课题为中心的分析过程:从数据中发现候选论点,通过跨源迭代验证,并发展为完整分析叙事。在 InsightBench 上,DataSTORM 实现新基准表现,洞察级召回率相对提升 19.4%,摘要级得分提升 7.2%。我们还基于真实复杂数据库 ACLED 构建新数据集,结果表明,DataSTORM 在自动化指标与人工评估上均优于 ChatGPT Deep Research 等商用系统。
原文摘要 · Abstract (English)
Deep research with Large Language Model (LLM) agents is emerging as a powerful paradigm for multi-step information discovery, synthesis, and analysis. However, existing approaches primarily focus on unstructured web data, while the challenges of conducting deep research over large-scale structured databases remain relatively underexplored. Unlike web-based research, effective data-centric research requires more than retrieval and summarization and demands iterative hypothesis generation, quantitative reasoning over structured schemas, and convergence toward a coherent analytical narrative. In this paper, we present DataSTORM, an LLM-based agentic system capable of autonomously conducting research across both large-scale structured databases and internet sources. Grounded in principles from Exploratory Data Analysis and Data Storytelling, DataSTORM reframes deep research over structured data as a thesis-driven analytical process: discovering candidate theses from data, validating them through iterative cross-source investigation, and developing them into coherent analytical narratives. We evaluate DataSTORM on InsightBench, where it achieves a new state-of-the-art result with a 19.4% relative improvement in insight-level recall and 7.2% in summary-level score. We further introduce a new dataset built on ACLED, a real-world complex database, and demonstrate that DataSTORM outperforms proprietary systems such as ChatGPT Deep Research across both automated metrics and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。