arXiv:2608.24825cs.AI2026-08

用大模型分析题库中隐含内容重复,提升测验公平性

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

论文配图:A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
图 1 · 摘自论文原文
  • 分结构与语义双维度评估题目相似性
  • 比传统方法更准识别无关内容冗余
  • 适合题库建设与自适应测验优化

大规模测评和自动命题的快速发展引发了对隐含内容冗余的担忧,即题干表述或情境设置等与测验构念无关的元素在题目间无意重复。传统相似度指标如BLEU或余弦相似度难以同时捕捉驱动感知冗余的结构与语义层面。本文提出一种基于大语言模型(LLM)的自动化题目相似性分析框架(AISA),通过结构分解与语义相关性双重机制实现相似性建模。心理测量验证表明,基于LLM的指标与无关局部依赖性指标更吻合,且能生成更合理的题目参数聚类。在计算机化自适应测试(CAT)中的模拟结果显示,引入基于LLM的相似性约束可提升参数估计稳定性并降低偏差,效率损失极小,优于传统指标约束。研究证明,该框架在支持可扩展题库管理、内容感知组卷及体验敏感型自适应测验方面具有潜力。

原文摘要 · Abstract (English)

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.

题库分析大模型自适应测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。