arXiv:2603.04670cs.AIcs.CL2026-03

用大模型分析图文结合的可视化题目,预测用户答对率。

Using Vision + Language Models to Predict Item Difficulty

  • 结合文本与图像特征,用大模型预测题目难度
  • 多模态方法误差最低,MAE仅0.224
  • 适合教育测评自动化与题目设计

本研究探究大型语言模型(LLMs)在判断数据可视化素养测试题难度方面的能力。我们考察仅使用题目文本(问题与选项)、仅使用可视化图像,或两者结合的特征,能否预测美国成人对题目的正确作答比例。采用GPT-4.1-nano分析不同特征集并生成难度预测。多模态方法(融合视觉与文本)取得最低平均绝对误差(MAE)0.224,优于仅用视觉(0.282)或仅用文本(0.338)的方法。最佳多模态模型在独立测试集上表现良好,均方误差达0.10805,展现了大模型在心理测量学分析与自动题目开发中的潜力。

原文摘要 · Abstract (English)

This project investigates the capabilities of large language models (LLMs) to determine the difficulty of data visualization literacy test items. We explore whether features derived from item text (question and answer options), the visualization image, or a combination of both can predict item difficulty (proportion of correct responses) for U.S. adults. We use GPT-4.1-nano to analyze items and generate predictions based on these distinct feature sets. The multimodal approach, using both visual and text features, yields the lowest mean absolute error (MAE) (0.224), outperforming the unimodal vision-only (0.282) and text-only (0.338) approaches. The best-performing multimodal model was applied to a held-out test set for external evaluation and achieved a mean squared error of 0.10805, demonstrating the potential of LLMs for psychometric analysis and automated item development.

大模型难度预测多模态教育测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。