新基准测试揭示大模型在多轮数据分析中表现不佳
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
- 用模拟用户逐步提问的方式评估大模型的多轮交互能力
- 顶级编码模型在不到50%的任务中达成正确结果
- 适合关注大模型分析可靠性与迭代推理的研究者
大语言模型(LLMs)在数据分析师角色中展现出潜力,但现有基准忽略了该领域中专家决策随数据洞察深化而迭代的特性。为此,我们提出IDA-Bench,一个新型基准,用于评估LLM代理在多轮交互场景下的表现。任务源自复杂的Kaggle笔记本,以逐轮自然语言指令形式呈现,由模拟用户的LLM生成。代理性能通过其最终数值输出与人工基准的对比来评估。初步结果显示,即使最先进的编码代理(如Claude-3.7-thinking)在少于50%的任务中成功,暴露出单轮测试中未显现的局限性。该研究强调提升LLM多轮能力的必要性,以构建更可靠的分析代理,需在指令遵循与推理之间取得平衡。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper insights of the dataset. To address this, we introduce IDA-Bench, a novel benchmark evaluating LLM agents in multi-round interactive scenarios. Derived from complex Kaggle notebooks, tasks are presented as sequential natural language instructions by an LLM-simulated user. Agent performance is judged by comparing its final numerical output to the human-derived baseline. Initial results show that even state-of-the-art coding agents (like Claude-3.7-thinking) succeed on < 50% of the tasks, highlighting limitations not evident in single-turn tests. This work underscores the need to improve LLMs' multi-round capabilities for building more reliable data analysis agents, highlighting the necessity of achieving a balance between instruction following and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。