用微调版ChatGPT自动评分中文科学解释,效果受推理复杂度影响。
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
- 微调ChatGPT适配中文科学解释评分任务
- 低阶回答越简洁越准,高阶回答越全面越准
- 适合教育评估中需自动化评分的场景
科学解释能力是科学素养的重要组成部分,但学生书面解释的评分仍具挑战且耗时。大型语言模型(LLMs)在英语等拼音语言中展现出自动评分潜力,但在表意文字如中文中的应用尚不明确。本研究探讨微调后的ChatGPT在中文科学解释自动评分中的表现。收集了学生对七个科学解释任务的作答,并通过肯德尔相关性分析评分准确率与推理复杂度的关系。定性分析揭示语言特征对评分准确性的影响。结果表明,领域适配使ChatGPT能有效评分中文科学解释。然而,评分准确率与推理复杂度呈负相关(低阶回答)和正相关(高阶回答)。模型对低阶回答中结构复杂的句子过度打分,而对高阶回答中简明因果表述则打分偏低。这种差异源于语言特征:低阶回答的简洁清晰提升准确率,高阶回答的信息全面性更关键。简单短句在低阶中更易得高分,长篇信息密集文本在高阶中更优。研究证实了大模型在中文教育评估中的有效性,并强调语言特征与推理复杂度在微调评分模型中的重要性。
原文摘要 · Abstract (English)
The development of explanations for scientific phenomena is essential in science assessment, but scoring student-written explanations remains challenging and resource-intensive. Large language models (LLMs) have shown promise in addressing this issue, particularly in alphabetic languages like English. However, their applicability to logographic languages is less explored. This study investigates the potential of fine-tuning ChatGPT, a leading LLM, to automatically score scientific explanations written in Chinese. Student responses to seven scientific explanation tasks were collected and automatically scored, with scoring accuracy examined in relation to reasoning complexity using the Kendall correlation. A qualitative analysis explored how linguistic features influenced scoring accuracy. The results show that domain-specific adaptation enables ChatGPT to score Chinese scientific explanations with accuracy. However, scoring accuracy correlates with reasoning complexity: a negative correlation for lower-level responses and a positive one for higher-level responses. The model overrates complex reasoning in low-level responses with intricate sentence structures and underrates high-level responses using concise causal reasoning. These correlations stem from linguistic features--simplicity and clarity enhance accuracy for lower-level responses, while comprehensiveness improves accuracy for higher-level ones. Simpler, shorter responses tend to score more accurately at lower levels, whereas longer, information-rich responses yield better accuracy at higher levels. These findings demonstrate the effectiveness of LLMs in automatic scoring within a Chinese context and emphasize the importance of linguistic features and reasoning complexity in fine-tuning scoring models for educational assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。