arXiv:2507.13881cs.CLcs.AI2025-07中稿 · presentation at th…被引 3

用大模型自动识别情境判断题中的个人与职业能力特征

Using LLMs to identify features of personal and professional skills in an open-response situational judgment test

  • 用大模型从开放回答中提取能力相关特征
  • 在Casper情境测试上验证了方法的有效性
  • 为自动化评估职业能力提供新思路,适合教育评估研究者

学术项目日益重视个人与职业能力,认为其与技术专长同样关键,有助于学生应对多样化的职业发展。随着这一需求增长,亟需可扩展的系统来测量、评估和发展这些能力。情境判断测试(SJTs)为标准化、可靠地评估这些能力提供了可能,但传统的开放式SJTs依赖受过训练的人类评分员,难以大规模实施。以往基于NLP的评分系统因构念效度不足而效果不佳。本文探索了一种利用大型语言模型(LLMs)从SJTs回答中提取与构念相关特征的新方法,并以Casper SJT为例验证了该方法的有效性。本研究为未来自动化评估个人与职业能力奠定了基础。

原文摘要 · Abstract (English)

Academic programs are increasingly recognizing the importance of personal and professional skills and their critical role alongside technical expertise in preparing students for future success in diverse career paths. With this growing demand comes the need for scalable systems to measure, evaluate, and develop these skills. Situational Judgment Tests (SJTs) offer one potential avenue for measuring these skills in a standardized and reliable way, but open-response SJTs have traditionally relied on trained human raters for evaluation, presenting operational challenges to delivering SJTs at scale. Past attempts at developing NLP-based scoring systems for SJTs have fallen short due to issues with construct validity of these systems. In this article, we explore a novel approach to extracting construct-relevant features from SJT responses using large language models (LLMs). We use the Casper SJT to demonstrate the efficacy of this approach. This study sets the foundation for future developments in automated scoring for personal and professional skills.

大模型能力评估情境测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。