构建首个面向交互式表格处理的综合评测基准,揭示大模型在复杂场景下的能力短板。
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs

- 设计四类18项任务,覆盖表格匹配、清洗、增强与转换,含1296个跨领域实例
- 在线评估模拟多轮交互与规则依赖,发现模型在复杂表和噪声对话中性能显著下降
- 适合研究交互式数据智能、评估大模型实际应用能力的研究者使用
表格数据处理是数据工作核心,基于大语言模型的助手虽展现出潜力,但现有评测多集中于单轮、完整指令下的表格推理,难以反映真实中多轮交互、需求演进的复杂处理过程。为此,我们提出CITBench,一个全面评估大模型在交互式表格处理中表现的基准。该基准涵盖表格匹配、清洗、增强和转换四大类别,包含18种任务类型和1296个实例,数据来自多个领域。支持离线与在线两种评估模式,其中在线模式模拟受限操作流程与结构化任务脚本下的多轮交互,捕捉用户参与式处理的关键行为特征。我们在CITBench上评估了大量开源与闭源大模型,结果表明:当前模型在简单表格和规则下表现良好,但随着表格复杂度提升、规则依赖变紧以及多轮交互噪声增加,性能显著下降。这揭示了大模型在理解、规划与表格结构感知方面在长期交互场景中的持续挑战。
原文摘要 · Abstract (English)
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primarily focus on table reasoning under single-turn, fully specified instructions, underrepresenting complex table processing that unfolds through multi-turn interactions with evolving user requirements. To bridge this gap, we introduce CITBench, a comprehensive benchmark for evaluating LLMs on interactive tabular data processing. CITBench features a comprehensive taxonomy across four high-level categories--table matching, cleaning, augmentation, and transformation--spanning 18 task types and 1,296 instances curated from datasets across diverse domains. The benchmark supports both offline and online evaluation, where the online setting models multi-turn interactions under constrained operation procedures and structured task scripts, capturing key potential behavioral characteristics of user-in-the-loop tabular data processing. We evaluate a broad suite of open-source and closed-source LLMs on CITBench, revealing a consistent trend: while current models perform well on simple tables and rules, their performance degrades significantly with increasing table complexity, tighter rule dependencies, and noisy multi-turn interaction simulations. These results highlight persistent challenges in understanding, planning, and table-structure awareness for LLMs in extended interactive data processing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。