arXiv:2608.10154cs.CLcs.AI2026-08中稿 · AIME-Con 2026

用多模态大模型拟合题目作答概率,直接预测题目难度。

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

  • 用微调的多模态LLM学习图文题目的作答概率模式。
  • 在保留能力水平标签的数据上,准确还原3PL和MCM模型曲线。
  • 适合教育测评、智能出题系统研究者参考。

我们展示了使用基于Qwen3.5微调的多模态大语言模型,重建多项选择模型(MCM)和三参数逻辑模型(3PL)曲线的结果。该模型通过提示与微调,在包含图文刺激的大规模多项选择题语料库上,学习不同学生能力水平下的选项选择概率。通过捕捉学生在离散能力区间上的系统性错误模式,模型隐式地学习到3PL和MCM曲线所编码的响应概率。这使得我们能够直接从模型预测的选项概率中,对预留测试集的题目难度进行高精度估计。

原文摘要 · Abstract (English)

We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a large training corpus of multiple-choice items containing both image and text stimuli, conditioned on a labeled set of student ability levels. By learning to reproduce the systematic error patterns of students across a discrete range of abilities, the LLM implicitly captures the underlying response probabilities encoded in the 3PL and MCM curves. This allows us to accurately approximate item difficulty on a held-out test set directly from the model's predicted option probabilities.

多模态教育测评大模型题目难度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。