arXiv:2506.19794cs.CLcs.AI2025-06AAAI被引 5

开源大模型分析数据能力弱?研究发现规划能力是关键。

Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study

  • 通过真实场景数据集评估模型在理解、编码和规划三方面表现
  • 规划质量决定模型表现,任务复杂度与交互设计影响推理能力
  • 数据质量比多样性更重要,新合成方法显著提升分析能力

大型语言模型在自动化数据分析任务中前景广阔,但开源模型在需要深度推理的场景中仍存在明显局限。本文通过构建涵盖多样真实场景的种子数据集,从数据理解、代码生成和策略规划三个核心维度评估模型表现。研究发现:(1) 策略规划质量是决定模型性能的关键因素;(2) 交互设计与任务复杂度显著影响推理能力;(3) 数据质量对性能提升的影响大于数据多样性。基于这些发现,我们提出一种数据合成方法,显著提升了开源LLM的数据分析推理能力。代码已公开于 https://github.com/zjunlp/DataMind。

原文摘要 · Abstract (English)

Large Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate model behavior across three core dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Code is available at https://github.com/zjunlp/DataMind.

大模型数据分析推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。