arXiv:2503.02532cs.CYcs.CL2025-03中稿 · Publication in Edu…被引 4

用AI自动评估提示词能力,让学习更高效。

Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development

  • 基于少量标注样本和规则生成特征,构建少样本检测器。
  • GPT-4在多数特征识别上表现优异,其他模型结果不一。
  • 适合教育研究者与AI学习工具开发者参考。

大型语言模型(LLM)驱动的聊天机器人(如ChatGPT)已在多个领域广泛应用,支持各类任务流程。然而,由于LLM内在复杂性,有效提示设计远比表面困难。这凸显了开发既普及又无缝集成于工作流中的教学与支持策略的必要性。但提示技能高度依赖任务与领域,通用方法效果有限。本研究探索是否可通过基于LLM的方法,利用临时指南和极少标注提示样本来实现学习评估。框架将指南转化为可识别于学习者提示中的特征,并结合标注样例构建少样本学习检测器。我们测试三种主流LLM及集成模型的不同配置,在原始提示样本上进行交叉验证,并对无任务经验学习者的提示进行测试。结果显示,GPT-4在多数特征检测中表现良好,而同类模型如GPT-3与GPT-3.5 Turbo(Instruct)在特征分类上表现出不一致行为。这些差异凸显了设计选择对特征筛选与提示识别影响的进一步研究需求。研究成果为生成式AI素养与计算机支持学习评估提供重要洞见,对研究者与实践者均有价值。

原文摘要 · Abstract (English)

The use of large language model (LLM)-powered chatbots, such as ChatGPT, has become popular across various domains, supporting a range of tasks and processes. However, due to the intrinsic complexity of LLMs, effective prompting is more challenging than it may seem. This highlights the need for innovative educational and support strategies that are both widely accessible and seamlessly integrated into task workflows. Yet, LLM prompting is highly task- and domain-dependent, limiting the effectiveness of generic approaches. In this study, we explore whether LLM-based methods can facilitate learning assessments by using ad-hoc guidelines and a minimal number of annotated prompt samples. Our framework transforms these guidelines into features that can be identified within learners' prompts. Using these feature descriptions and annotated examples, we create few-shot learning detectors. We then evaluate different configurations of these detectors, testing three state-of-the-art LLMs and ensembles. We run experiments with cross-validation on a sample of original prompts, as well as tests on prompts collected from task-naive learners. Our results show how LLMs perform on feature detection. Notably, GPT- 4 demonstrates strong performance on most features, while closely related models, such as GPT-3 and GPT-3.5 Turbo (Instruct), show inconsistent behaviors in feature classification. These differences highlight the need for further research into how design choices impact feature selection and prompt detection. Our findings contribute to the fields of generative AI literacy and computer-supported learning assessment, offering valuable insights for both researchers and practitioners.

提示工程AI评估少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。