arXiv:2608.04554cs.CLcs.CV2026-08

对比图文表述与图像原生建模,提升数学题难度预测准确率

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

论文配图:Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
图 1 · 摘自论文原文
  • 用语言描述图像或保留原始图像两种方式建模视觉信息
  • 图像原生建模在多数模型上表现更优,尤其经广泛适配后
  • 两种方法错误互补,适合不同技术场景的开发者参考

从内容预测题目难度可在学生作答数据不足时提供初始估计。现有方法通常将题干和选项表示为文本,对含视觉成分的数学题,常先将视觉证据转为语言描述,再用文本模型预测。我们探讨:视觉证据应如何表示以提升难度预测?基于来自Eedi、经学生作答校准难度的题目,我们直接使用大语言模型(LLMs)和视觉语言模型(VLMs)进行难度回归训练。结果表明,两种视觉处理方式均优于仅用文本,其中开放型视觉语言模型(Open-VLM)在所有测试的LLM中取得更低的均方根误差(RMSE),而广泛适配的图像原生模型在所有评估的VLM中表现更优。测试时干预显示,模型仍依赖完整题目图像,但无法单独分离视觉部分贡献。两种方法产生部分互补的个体错误,且计算流程差异显著。因此,图文表述不应被视为唯一可行方案,图像原生建模是同样有效的替代选择,其效果取决于模型适配程度。

原文摘要 · Abstract (English)

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textualizes that evidence and then applies a text predictor. We ask: how should visual evidence be represented for item difficulty prediction? We compare question text alone, visual textualization, which expresses visual evidence in language, and image-native modeling, which retains the original image. Using Eedi items with difficulty calibrated from student responses, we train large language models (LLMs) and vision-language models (VLMs) directly for difficulty regression. Both visual interfaces achieve the lowest point estimates, although the leading systems cannot be reliably ordered. Open-VLM textualization yields lower RMSE point estimates for all evaluated LLMs, while broader adaptation does so for all image-native VLMs. Test-time interventions show dependence on the paired full-item image, but do not isolate the additional visual component. The two visual interfaces also make partially complementary item-level errors and differ substantially in computational workflow. Thus, textualization should not be treated as the only practical interface: image-native modeling is a competitive alternative whose effectiveness depends on how the VLM is adapted.

难度预测视觉建模VLM教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。