LLM自动化叙事存疑,实测显示其表现不如人类专家稳定可靠。
Flaws in the LLM Automation Narrative
- 设计新代码任务评估人类与大模型的分析能力差异。
- 人类专家平均表现更优,且结果波动更小、错误更少。
- 强调评估时需关注误差大小和结果方差,而非仅看平均分。
大型语言模型(LLMs)被广泛认为在知识经济任务中达到人类专家水平,这一判断主要基于其在标准化数据集上的平均表现基准测试。然而,多数基准测试存在局限:它们常依赖模型训练数据中的直接内容,且未评估性能可靠性或误差幅度。在高风险场景下,这些特性至关重要。本文通过一项新型基准任务——编写代码完成数据分析,对比前沿大模型与人类专家的表现,并显式测量响应的方差与误差大小。结果显示,人类专家在多项指标上平均表现更优,且结果变异更小。研究证明,大模型并未持续达到人类专家水平,强调在评估中必须纳入方差和误差幅度分析。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly described as performing at the level of human experts on knowledge economy tasks. These claims are primarily based on how LLMs perform on benchmarking tasks that measure average performance across standardized datasets. Primary limitations of many benchmarking tasks are that they often measure performance based on content directly included in LLM training data, and they frequently do not assess the reliability of LLM performance or the magnitude of LLM errors. However, in high stakes contexts, these qualities are critically important. Through a novel LLM benchmarking task that requires writing computer code to complete a data analysis task, we compare the performance of a frontier LLM against submissions from human experts and explicitly measure the variance of responses and the magnitude of errors. Our study reveals that the human experts perform better on average on a range of metrics and demonstrate less variability in performance. Our results provide evidence that LLMs do not consistently perform at the level of human experts and demonstrate the importance of measuring variance and assessing error magnitude in LLM benchmark evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。