调优推理策略比模型集成更能提升数学对话评分准确率与性价比。
The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost
- 用自一致性投票和调整推理强度优化评分,效果优于多模型集成。
- 推理力度越高,评分准确率越强,且呈线性增长趋势。
- Gemini 3.1 Pro 低推理配置最准但贵,GPT-5.4 Nano 无推理最划算。
为优化大语言模型在高中数学对话题上的自动评分,我们考察了自一致性(模型内多数投票)与推理努力程度的影响。基于 OpenAI 与 Google 的前沿与低成本模型,评估了 900 条学生对话与人工标注的基准分数。温度采样显著提升准确率,但增加集成规模(j=1 到 7)未带来显著收益。推理努力程度与评分准确率呈显著正向线性关系,且不同模型家族表现差异明显。效率前沿分析显示:Gemini 3.1 Pro Preview 低推理配置最准确但成本高;而 GPT-5.4 Nano 与 Mini 无推理配置在成本与性能间达到最佳平衡。
原文摘要 · Abstract (English)
Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs). We examined self-consistency (intra-model majority voting) and reasoning effort for scoring conversation-based assessment items in high school mathematics, evaluating 900 student conversations against human-scored ground truths using frontier and low-cost models from OpenAI and Google. Temperature sampling significantly improved accuracy over deterministic calls, but increasing ensemble size (j = 1 to 7) produced no significant gains. Higher reasoning effort showed a significant positive linear trend with scoring accuracy, though the benefit varied by model family. An efficiency frontier analysis identified Gemini 3.1 Pro Preview at low reasoning as the most accurate but costly configuration; GPT-5.4 Nano and Mini with no reasoning offered the best cost-performance balance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。