用多种方法分析数学分班考,发现题6最有效,建议改两阶段考试。
Multi-Method Analysis of Mathematics Placement Assessments: Classical, Machine Learning, and Clustering Approaches
- 融合经典理论、机器学习与聚类,多角度评估考题质量。
- 题6区分度达1.0,贡献20.6%预测力,是核心题。
- 发现应改55%门槛为42.5%,避免误判学生能力。
本研究对198名学生参加的40题数学分班考试,采用经典测验理论、机器学习和无监督聚类的多方法框架进行评估。经典测验理论显示,55%题目具备优秀区分度(D≥0.40),但30%题目区分度差(D<0.20)需替换。第6题(图像理解)表现最优,区分度达1.000,方差分析F值最高(F=4609.1),随机森林特征重要性达0.206,贡献20.6%预测力。机器学习模型表现优异,随机森林与梯度提升交叉验证准确率分别达97.5%和96.0%。K均值聚类识别出天然二分能力结构,临界点为42.5%,低于机构现行55%标准,暗示过度分类风险。双聚类方案稳定性高(引导自举ARI=0.855),低分群纯度完美。多方法结果一致支持:替换低区分度题目、实施两阶段评估、结合随机森林预测并保障透明性。结果表明,多方法整合为数学分班优化提供稳健实证基础。
原文摘要 · Abstract (English)
This study evaluates a 40-item mathematics placement examination administered to 198 students using a multi-method framework combining Classical Test Theory, machine learning, and unsupervised clustering. Classical Test Theory analysis reveals that 55\% of items achieve excellent discrimination ($D \geq 0.40$) while 30\% demonstrate poor discrimination ($D < 0.20$) requiring replacement. Question 6 (Graph Interpretation) emerges as the examination's most powerful discriminator, achieving perfect discrimination ($D = 1.000$), highest ANOVA F-statistic ($F = 4609.1$), and maximum Random Forest feature importance (0.206), accounting for 20.6\% of predictive power. Machine learning algorithms demonstrate exceptional performance, with Random Forest and Gradient Boosting achieving 97.5\% and 96.0\% cross-validation accuracy. K-means clustering identifies a natural binary competency structure with a boundary at 42.5\%, diverging from the institutional threshold of 55\% and suggesting potential overclassification into remedial categories. The two-cluster solution exhibits exceptional stability (bootstrap ARI = 0.855) with perfect lower-cluster purity. Convergent evidence across methods supports specific refinements: replace poorly discriminating items, implement a two-stage assessment, and integrate Random Forest predictions with transparency mechanisms. These findings demonstrate that multi-method integration provides a robust empirical foundation for evidence-based mathematics placement optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。