arXiv:2602.17377cs.CL2026-02

发现正确选项在语料库中出现更频繁,可凭此提升选择题答题准确率。

Corpus Prevalence of Multiple-Choice Question Options

  • 用文本嵌入分析大规模语料库中选项的出现频率,识别正确答案
  • 选最常见选项可比随机猜测高出9.0%准确率
  • AI生成干扰项也呈现类似频率模式,适合教育类AI系统优化

近年来,基于语料库的AI方法(如大语言模型)在教育领域广泛应用。然而,其统计性本质可能影响行为表现。本文聚焦于多选题(MCQ)中选项的语料库频率问题,探究:仅凭选项在语料中的出现频率是否足以判断正确答案?我们提出一种基于文本嵌入的计算方法,评估专家与机器生成的多个大型题集中的选项频率。结果表明,在三个大规模题集中,正确答案无论题干如何,均显著比错误选项更常见。以维基百科为检索语料库,仅选择最频繁选项即可达到比随机猜测高出9.0%的准确率。此外,大模型生成的干扰项也表现出与人工设计相似的频率模式,尽管其训练数据庞大且具统计性。值得注意的是,语料频率并不等同于人类识别度。这揭示了需深入理解语料在教育类AI应用中的作用,无论直接使用还是通过大模型间接实现。

原文摘要 · Abstract (English)

In recent years, corpus-driven AI methods, such as Large Language Models (LLMs), have seen widespread use in education. While on the surface their abilities look promising for tasks ranging from generating assessment materials to simulating student performance, we should be aware of the subtle nuances of their frequentist nature that might be affecting their behaviour. In this work, we focus on the aspect of corpus frequency in the context of creating high-quality Multiple Choice Questions (MCQs), specifically asking: What if corpus prevalence were enough to identify the correct answer to an MCQ? We propose a computational method of assessing corpus prevalence of MCQ options in large text corpora leveraging textual embeddings using both expert- and machine-generated MCQ sets. The key finding, across three large question sets, is that correct answers, independently of the question stem, are significantly more available than incorrect options. Specifically, using Wikipedia as the retrieval corpus, we find that always selecting the most prevalent option leads to scores up to 9.0% above the random-guess baseline. We also find that MCQ distractors generated by LLMs often show similar patterns of prevalence compared to expert-created options, despite the LLMs' frequentist nature and their training on large collections of textual data. Moreover, we find that corpus prevalence does not necessarily correlate with how recognisable terms are to humans. This highlights the need to better understand how corpora are used in AI-driven methods for education, whether applied directly or indirectly via LLMs.

多选题语料频率大模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。