开源大模型在作文评分中表现媲美闭源模型,成本更低且更公平。
Bridging the LLM Accessibility Divide? Performance, Fairness, and Cost of Closed versus Open LLMs for Automated Essay Scoring
- 对比九个主流大模型,测试其在作文评分与生成中的表现。
- 开源模型如Llama 3与GPT-4在评分准确率上无显著差异,成本低37倍。
- 适合关注模型公平性、成本控制与可及性的教育AI研究者使用。
闭源大语言模型(如GPT-4)在多项自然语言处理任务中达到领先水平,但其可用性、成本和透明度引发广泛争议。本研究对九个主流大模型(涵盖闭源、开源与开放源代码生态)在自动作文评分相关的文本评估与生成任务中进行严谨对比。结果表明,在基于少样本学习的人类作文评估中,开源模型如Llama 3与Qwen2.5在预测性能上与GPT-4相当,年龄或种族相关偏差影响无显著差异;其中Llama 3的成本效率比GPT-4高至37倍。在生成任务中,顶级开源模型生成的作文在语义构成、嵌入表示及机器评估得分上也与闭源模型相当。研究挑战了闭源模型的主导地位,凸显开源模型在提升可及性方面的潜力,表明其可在保持竞争力的同时缩小技术鸿沟。
原文摘要 · Abstract (English)
Closed large language models (LLMs) such as GPT-4 have set state-of-the-art results across a number of NLP tasks and have become central to NLP and machine learning (ML)-driven solutions. Closed LLMs' performance and wide adoption has sparked considerable debate about their accessibility in terms of availability, cost, and transparency. In this study, we perform a rigorous comparative analysis of nine leading LLMs, spanning closed, open, and open-source LLM ecosystems, across text assessment and generation tasks related to automated essay scoring. Our findings reveal that for few-shot learning-based assessment of human generated essays, open LLMs such as Llama 3 and Qwen2.5 perform comparably to GPT-4 in terms of predictive performance, with no significant differences in disparate impact scores when considering age- or race-related fairness. Moreover, Llama 3 offers a substantial cost advantage, being up to 37 times more cost-efficient than GPT-4. For generative tasks, we find that essays generated by top open LLMs are comparable to closed LLMs in terms of their semantic composition/embeddings and ML assessed scores. Our findings challenge the dominance of closed LLMs and highlight the democratizing potential of open LLMs, suggesting they can effectively bridge accessibility divides while maintaining competitive performance and fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。