GPT-4o-mini可稳定评分音乐分析作文,但不同提示策略效果差异大。
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

- 用三种提示策略对300篇学生作答评分,对比与教师均分的一致性。
- 少样本+思维链策略最接近教师评分,检索增强法系统偏高,自一致性重复性好但精度低。
- 术语部分评分一致性较差,需逐维度验证,仍需人工把关。
评分开放式的音乐分析作答耗时且依赖对和声与结构的精细判断。本研究评估了 GPT-4o-mini 在基于评分量规下对音乐分析论文进行自动评分的有效性与可重复性,以教师均分为基准。使用包含300篇大学水平学生作答的数据集,由教师在四个维度(和声、结构、推理、术语)上打分。GPT-4o-mini 采用三种提示策略:少样本+思维链(Fs+CoT)、检索增强生成(RAG)和基于五次内部生成的自一致性(SC),每种策略运行三次,模型、提示、量规与作答保持不变。单次评分模拟实际操作,中位数聚合用于检验鲁棒性。通过相关系数、组内相关系数、Krippendorff's alpha、加权卡帕系数及评分误差指数评估与教师均分的一致性。结果显示,Fs+CoT在单次评分与中位数聚合中均与教师均分一致度最高;RAG存在系统性高估;SC得分高度可重复但个体层面一致性较弱。维度分析表明,各量规维度表现不一,术语维度一致性普遍低于推理。结果表明,GPT-4o-mini 可对复杂音乐分析作答生成稳定评分,但不同提示策略产生不同的评分特征。因此,实际应用需进行策略特异性校准、维度级验证并持续保留人工监督。
原文摘要 · Abstract (English)
Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff's alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。