arXiv:2501.01588cs.CLcs.AI2025-01

微调小模型PHI-3,让其答题准确率从62%升至90.8%

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges

  • 用TruthfulQA数据集微调PHI-3,优化提示词设计
  • 微调后困惑度从4.68降至2.27,准确率提升至90.8%
  • 适合教育场景中的智能评测与个性化学习应用

大型语言模型在理解与生成类人文本方面表现出色,但在多选题回答任务中仍面临幻觉和提示不清晰等挑战。本文研究微软的紧凑高效模型PHI-3在多选题问答中的潜力。通过在TruthfulQA数据集上微调,并设计优化提示,结合困惑度、准确率与F1分数评估。结果显示,PHI-3.5微调后困惑度由4.68降至2.27,准确率从62%提升至90.8%。研究强调了高效模型在自适应学习系统与教育评估中的重要性,为课堂集成提供支持,尤其适用于考试准备、学生反馈与个性化学习。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become essential tools across various domains due to their impressive capabilities in understanding and generating human-like text. The ability to accurately answer multiple-choice questions (MCQs) holds significant value in education, particularly in automated tutoring systems and assessment platforms. However, adapting LLMs to handle MCQ tasks effectively remains challenging due to the hallucinations and unclear prompts. This work explores the potential of Microsoft's PHI-3\cite{Abdin2024}, a compact yet efficient LLM, for MCQ answering. Our contributions include fine-tuning the model on the TruthfulQA dataset, designing optimized prompts to enhance model performance, and evaluating using perplexity and traditional metrics like accuracy and F1 score. Results show a remarkable improvement in PHI-3.5's MCQ handling post-fine-tuning, with perplexity decreasing from 4.68 to 2.27, and accuracy rising from 62\% to 90.8\%. This research underlines the importance of efficient models in adaptive learning systems and educational assessments, paving the way for broader integration into the classroom, particularly in fields like test preparation, student feedback, and personalized learning.

多选题微调教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。