arXiv:2508.19014cs.AI2025-08

不依赖语言分析,仅用答题得分和用时估算数学题难度。

MAB Optimizer for Estimating Math Question Difficulty via Inverse CV without NLP

  • 用强化学习的多臂老虎机模型,从解题数据中自动估计题目难度。
  • 在三个数据集上平均R²达0.9213,均方根误差仅0.0584。
  • 适合教育系统自适应评估,尤其适用于代数等符号领域。

技术与教育的融合催生了智能自主辅导系统(IATS),客观且领域无关的题目难度评估至关重要。传统人工标注主观性强,现有基于自然语言处理的方法在代数等符号领域表现不佳。本研究提出被动学习者测量法(APME),一种基于强化学习的多臂老虎机(MAB)框架,仅利用解题得分与耗时数据,无需语言特征或专家标签即可估算难度。通过逆变异系数作为风险调整指标,模型提供可解释且可扩展的自适应评估机制。在三个异构数据集上的实证验证显示,模型平均R²达0.9213,平均RMSE为0.0584,证明其鲁棒性、准确性和跨教育阶段与评估形式的适应能力。相比回归、NLP驱动及项目反应理论(IRT)等基线方法,该框架在纯符号领域表现更优。研究发现:(i)题目异质性显著影响感知难度;(ii)解题结果方差与均值同等重要,对自适应分配具有关键作用。从教学角度,模型契合维果茨基最近发展区理论,能识别挑战与可达成之间的平衡任务,维持学习动机并减少倦怠。该领域无关、自监督方法推动了IATS中的难度标注发展,可推广至任何存在解题交互数据的场景。

原文摘要 · Abstract (English)

The evolution of technology and education is driving the emergence of Intelligent & Autonomous Tutoring Systems (IATS), where objective and domain-agnostic methods for determining question difficulty are essential. Traditional human labeling is subjective, and existing NLP-based approaches fail in symbolic domains like algebra. This study introduces the Approach of Passive Measures among Educands (APME), a reinforcement learning-based Multi-Armed Bandit (MAB) framework that estimates difficulty solely from solver performance data -- marks obtained and time taken -- without requiring linguistic features or expert labels. By leveraging the inverse coefficient of variation as a risk-adjusted metric, the model provides an explainable and scalable mechanism for adaptive assessment. Empirical validation was conducted on three heterogeneous datasets. Across these diverse contexts, the model achieved an average R2 of 0.9213 and an average RMSE of 0.0584, confirming its robustness, accuracy, and adaptability to different educational levels and assessment formats. Compared with baseline approaches-such as regression-based, NLP-driven, and IRT models-the proposed framework consistently outperformed alternatives, particularly in purely symbolic domains. The findings highlight that (i) item heterogeneity strongly influences perceived difficulty, and (ii) variance in solver outcomes is as critical as mean performance for adaptive allocation. Pedagogically, the model aligns with Vygotskys Zone of Proximal Development by identifying tasks that balance challenge and attainability, supporting motivation while minimizing disengagement. This domain-agnostic, self-supervised approach advances difficulty tagging in IATS and can be extended beyond algebra wherever solver interaction data is available

难度评估自适应学习强化学习教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。