评测三款大模型在编程领域个性化职业辅导表现,发现GPT-4最精准贴心。
Assessing Personalized AI Mentoring with Large Language Models in the Computing Field
- 用零样本方式测试三款大模型对不同性别、种族、职级学生的定制化建议。
- GPT-4生成内容更贴合学生背景,关键词分析显示其个性化程度最高。
- 人类专家评价一致认可GPT-4表现最优,适合用于构建带人性温度的辅导系统。
本文深入评估了三款先进大语言模型(GPT-4、LLaMA 3、Palm 2)在计算领域个性化职业辅导中的表现,采用三种涵盖性别、种族和职业水平的学生画像进行测试。通过无监督零样本学习方法,在无需人工干预下运行自定义自然语言处理分析流程,识别响应中反映学生特征的关键词。结果表明,GPT-4在生成内容的独特性和针对性上优于其他两模型。定性分析显示,人类专家评审也支持该结论:GPT-4提供的建议更具准确性与实用性,并能以鼓励性语言应对特定挑战。本研究为构建融合真人导师的大模型个性化辅导工具提供了基础。
原文摘要 · Abstract (English)
This paper provides an in-depth evaluation of three state-of-the-art Large Language Models (LLMs) for personalized career mentoring in the computing field, using three distinct student profiles that consider gender, race, and professional levels. We evaluated the performance of GPT-4, LLaMA 3, and Palm 2 using a zero-shot learning approach without human intervention. A quantitative evaluation was conducted through a custom natural language processing analytics pipeline to highlight the uniqueness of the responses and to identify words reflecting each student's profile, including race, gender, or professional level. The analysis of frequently used words in the responses indicates that GPT-4 offers more personalized mentoring compared to the other two LLMs. Additionally, a qualitative evaluation was performed to see if human experts reached similar conclusions. The analysis of survey responses shows that GPT-4 outperformed the other two LLMs in delivering more accurate and useful mentoring while addressing specific challenges with encouragement languages. Our work establishes a foundation for developing personalized mentoring tools based on LLMs, incorporating human mentors in the process to deliver a more impactful and tailored mentoring experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。