用生成式AI评分需更全面的证据支持,避免盲目信任。
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
- 对比传统特征模型与生成式AI评分的差异
- 基于6-12年级作文数据验证评分有效性
- 为生成式AI评分提供可操作的证据框架
大型语言模型和生成式人工智能的快速发展,使其在高风险测评中的应用日益可能。将生成式AI用于开放回答题评分尤其吸引人,因其无需手工设计特征,且可能超越传统方法。本文旨在揭示特征模型与生成式AI在开放回答评分中的区别,提出支持生成式AI评分系统结果解释与使用的有效性证据收集最佳实践。我们比较了人工评分、基于特征的自然语言处理AI评分系统与生成式AI系统所需的有效性证据。由于缺乏透明性及一致性等独特问题,生成式AI所需的证据比特征模型更广泛。基于6-12年级学生独立议论文的大规模数据集,本文展示了不同类型评分系统的有效性证据收集过程,并突显了在构建有效性论据时面临的诸多复杂性与考量。
原文摘要 · Abstract (English)
The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more likely. Use of generative AI in the scoring of constructed responses is particularly appealing because it reduces the effort required for handcrafting features in traditional AI scoring and might even outperform those methods. The purpose of this paper is to highlight the differences in the feature-based and generative AI applications in constructed response scoring systems and propose a set of best practices for the collection of validity evidence to support the use and interpretation of constructed response scores from scoring systems using generative AI. We compare the validity evidence needed in scoring systems using human ratings, feature-based natural language processing AI scoring engines, and generative AI. The evidence needed in the generative AI context is more extensive than in the feature-based scoring context because of the lack of transparency and other concerns unique to generative AI such as consistency. Constructed response score data from a large corpus of independent argumentative essays written by 6-12th grade students demonstrate the collection of validity evidence for different types of scoring systems and highlight the numerous complexities and considerations when making a validity argument for these scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。