arXiv:2501.05891cs.CLcs.AI2025-01

小模型微调后答编程语言选择题更准,省钱省资源。

Affordably Fine-tuned LLMs Provide Better Answers to Course-specific MCQs

  • 用课程教材微调小模型,比大模型更高效
  • 7B模型经微调后准确率达82.1%,超越未微调的大模型
  • 适合教育场景快速部署,无需昂贵硬件

在教育领域,大语言模型生成类人文本的能力激发了其提升教与学效率的研究。本文研究在硬件限制下,不同精炼技术对回答多选题(MCQs)的影响。以编程语言课程的162道本科级题目为测试集(该数据集为本文贡献并公开),考察通用预训练模型(LLaMA-2的7B、13B、70B版本)在使用课程教材片段进行微调及量化压缩后的表现。结果表明,基于教材微调的小模型(如7B)在准确率上优于未微调的大模型,尤其在资源受限条件下更具优势,证明了教学场景中低成本高效使用大模型的可能性。

原文摘要 · Abstract (English)

In education, the capability of generating human-like text of Large Language Models (LLMs) inspired work on how they can increase the efficiency of learning and teaching. We study the affordability of these models for educators and students by investigating how LLMs answer multiple-choice questions (MCQs) with respect to hardware constraints and refinement techniques. We explore this space by using generic pre-trained LLMs (the 7B, 13B, and 70B variants of LLaMA-2) to answer 162 undergraduate-level MCQs from a course on Programming Languages (PL) -- the MCQ dataset is a contribution of this work, which we make publicly available. Specifically, we dissect how different factors, such as using readily-available material -- (parts of) the course's textbook -- for fine-tuning and quantisation (to decrease resource usage) can change the accuracy of the responses. The main takeaway is that smaller textbook-based fine-tuned models outperform generic larger ones (whose pre-training requires conspicuous resources), making the usage of LLMs for answering MCQs resource- and material-wise affordable.

大模型微调教育AI多选题生成资源优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。