arXiv:2510.13835cs.CLcs.AI2025-10被引 3

构建交互式数据分析评测基准,检验大模型长期协作能力

ConDABench: Interactive Evaluation of Language Models for Data Analysis

  • 基于论文生成真实对话数据解析任务,模拟用户意图澄清过程
  • 包含1420个对话分析问题,揭示新模型在长流程任务中表现仍不足
  • 适合评估大模型在复杂互动任务中的协作能力,推动人机协同发展

现实中的数据分析任务常存在目标不明确和数据脏乱的问题,需要通过用户交互来理解并澄清意图,因此交互是解决这类复杂任务的关键。现有评估大模型在数据分析任务上的基准无法体现这些复杂性,也缺乏对交互性的原生支持。我们提出ConDABench,一个用于生成对话式数据分析(ConDA)基准并评估外部工具的框架。该框架包括:(a) 从描述公共数据集洞察的文章中生成真实基准的多智能体工作流;(b) 使用该工作流生成的1,420个ConDA问题;(c) 首次实现系统化评估对话式数据分析工具在生成的ConDA问题上的评价工具包。对前沿大模型在该基准上的评估显示,尽管新一代模型能解决更多实例,但在需要持续、长周期互动的任务上并未表现出明显优势。ConDABench为模型开发者提供了一条衡量迈向真正协作型模型进展的路径。

原文摘要 · Abstract (English)

Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, essential to solving these complex tasks. Existing benchmarks for evaluating LLMs on data analysis tasks do not capture these complexities or provide first-class support for interactivity. We introduce ConDABench, a framework for generating conversational data analysis (ConDA) benchmarks and evaluating external tools on the generated benchmarks. \bench consists of (a) a multi-agent workflow for generating realistic benchmarks from articles describing insights gained from public datasets, (b) 1,420 ConDA problems generated using this workflow, and (c) an evaluation harness that, for the first time, makes it possible to systematically evaluate conversational data analysis tools on the generated ConDA problems. Evaluation of state-of-the-art LLMs on the benchmarks reveals that while the new generation of models are better at solving more instances, they are not necessarily better at solving tasks that require sustained, long-form engagement. ConDABench is an avenue for model builders to measure progress towards truly collaborative models that can complete complex interactive tasks.

对话分析交互评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。