arXiv:2511.22884cs.AI2025-11ACL被引 2

构建新基准InsightEval,评估大模型在数据中发现深层洞察的能力

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

  • 基于对现有基准缺陷的分析,设计高标准数据筛选流程
  • 提出新评测指标,量化智能体探索发现能力,揭示当前技术瓶颈
  • 适合关注大模型数据分析与自动洞察研究的学者与工程师

数据挖掘已成为科研不可或缺的环节。为发掘海量数据中隐藏的潜在知识与洞见,需开展深度探索性分析以释放其全部价值。随着大语言模型(LLMs)与多智能体系统的兴起,越来越多研究人员利用这些技术进行洞察发现。然而,目前缺乏有效的评估基准来衡量洞察发现能力。现有最全面的框架InsightBench仍存在格式不一致、目标设计不合理、洞察重复等严重缺陷,可能严重影响数据质量与智能体评估效果。为此,我们系统分析InsightBench的不足,提出高质量洞察评估基准的核心标准,并构建名为InsightEval的新数据集。我们进一步提出一种新型度量方法,用于评估智能体的探索性能。在InsightEval上开展大量实验,揭示了自动化洞察发现中的普遍挑战,并提出若干关键发现,为该前沿方向的未来研究提供指引。

原文摘要 · Abstract (English)

Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets, we need to perform deep exploratory analysis to realize their full value. With the advent of large language models (LLMs) and multi-agent systems, more and more researchers are making use of these technologies for insight discovery. However, there are few benchmarks for evaluating insight discovery capabilities. As one of the most comprehensive existing frameworks, InsightBench also suffers from many critical flaws: format inconsistencies, poorly conceived objectives, and redundant insights. These issues may significantly affect the quality of data and the evaluation of agents. To address these issues, we thoroughly investigate shortcomings in InsightBench and propose essential criteria for a high-quality insight benchmark. Regarding this, we develop a data-curation pipeline to construct a new dataset named InsightEval. We further introduce a novel metric to measure the exploratory performance of agents. Through extensive experiments on InsightEval, we highlight prevailing challenges in automated insight discovery and raise some key findings to guide future research in this promising direction.

大模型数据洞察评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。