用框选答题区域提升小模型阅卷准确率与效率
Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks

- 用边界框裁剪答题区域,聚焦关键信息
- 准确率显著提升,计算量降低超60%
- 适合大规模教育评估中部署小模型
将小型语言模型(SLMs)应用于教育场景可带来隐私保护、成本低和可扩展性强等优势。然而,由于处理大尺寸图像的计算开销高,且整页存在视觉干扰,SLMs在复杂视觉任务(如批改手写试卷)中表现不佳。本文研究通过边界框裁剪学生作答区域,能否提升SLMs在简答题批改任务中的准确率与计算效率。基于2025年澳大利亚物理奥林匹克竞赛的手写答卷扫描数据集,我们评估了4B至72B参数量的多个模型,在不同思维链(CoT)提示与图像裁剪条件下的表现。结果表明,使用边界框显著提升了各模型的评分准确率,并大幅降低计算成本(FLOPs)。结论是,边界框预处理是大规模视觉化教育评估中部署SLMs的关键步骤。
原文摘要 · Abstract (English)
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalability. However, SLMs often struggle with complex vision-based tasks, such as grading handwritten student exams, due to the high computational cost of processing large images and the visual distractions present on a full page. In this paper, we investigate whether cropping student responses using bounding boxes can improve the accuracy and computational efficiency of SLMs on a short-answer grading task. Using a dataset of scanned handwritten responses from the 2025 Australian Physics Olympiad, we evaluate the performance of several models ranging from 4B to 72B parameters under varying conditions of Chain of Thought (CoT) prompting and image cropping. Our results demonstrate that using bounding boxes significantly improves grading accuracy and reduces computational cost (FLOPs) across models. We conclude that bounding boxes are a crucial pre-processing step for deploying SLMs in large-scale, vision-based educational assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。