arXiv:2606.25491cs.CVcs.AI2026-06

构建手写作业多页答案区域定位基准,精准识别答案与推理步骤空间位置。

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

论文配图:HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment
图 1 · 摘自论文原文
  • 提出层级式答案区域标注,支持完整答案与步骤级子区域定位。
  • 零样本模型在整体定位和步骤分解上分别不超过55.22%和48.22%准确率。
  • 适用于教育智能评估、手写识别与视觉语言模型研究者。

自动作业评分不仅需识别学生作答内容,还需准确定位每道题及中间推理步骤在杂乱多页手写内容中的空间位置。本文针对缺乏评估标准的页面感知、两级答案区域定位任务,提出HG-Bench基准:从1,489,278张图像中精选500个K-12作业样本,包含问题级与步骤级标注框,并通过层级包含约束关联。该基准配套页面感知评估协议,分别衡量完整答案定位(FA)与步骤级分解(FSm),以检验模型是否真正理解学生推理的空间结构,而非仅解析可见文本。在前沿闭源API与主流开源视觉语言模型中,无零样本系统在FA上超过55.22%,在FSm上超过48.22%;而基于约1万条领域内数据微调的GLM-4.6V 9B模型达到74.97/72.26的性能。结果揭示了步骤级手写定位能力的显著差距,为未来自动作业评估研究提供可复现的基准、评估协议与参考模型。

原文摘要 · Abstract (English)

Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noisy, multi-page handwritten work. This paper addresses the missing evaluation setting of page-aware, two-level answer-region grounding: given a sequence of homework page images, a model must localize complete answer regions and their ordered step-level subregions. We introduce HG-Bench, a benchmark of 500 human-annotated K-12 homework samples curated from a 1,489,278-image source pool, with question-level and step-level boxes linked by a hierarchical containment constraint. HG-Bench is paired with a page-aware evaluation protocol that separately measures complete-answer localization (FA) and step-level decomposition (FSm), revealing whether models truly ground the spatial structure of student reasoning rather than merely parse visible text. Across frontier closed-source APIs and competitive open-weight VLMs, no zero-shot system exceeds 55.22% on FA or 48.22% on FSm, while a GLM-4.6V 9B reference model fine-tuned on ~10k in-domain examples reaches 74.97/72.26. These results identify step-level handwritten grounding as a concrete capability gap and provide a reproducible benchmark, evaluation protocol, and trained reference point for future work on automated homework assessment.

作业评估手写识别视觉语言模型多页定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。