用不确定性感知方法提升阿拉伯语可读性分级准确率
mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment
- 基于共形预测生成有覆盖率保证的预测集,结合软标签重归一化加权
- 在不同模型上稳定提升QWK 1-3分,句子级测试得分84.9%、盲测85.7%
- 适合需可解释分级结果的阿拉伯语教育评估场景
我们提出一种简单、模型无关的后处理技术,用于BAREC 2025共享任务中的细粒度阿拉伯语可读性分类(19个有序等级)。该方法应用共形预测生成具有覆盖率保证的预测集,并对共形集内的概率进行softmax重归一化后计算加权平均。这种不确定性感知解码有效减少了高惩罚误分类,使错误趋向于邻近等级。该方法在不同基础模型上均带来1-3点的QWK提升。在严格赛道中,我们的提交在句子级获得84.9%(测试集)和85.7%(盲测)的QWK得分,文档级为73.3%。对于阿拉伯语教育评估,该方法使人工评审员只需关注少数可能等级,兼具统计保证与实际可用性。
原文摘要 · Abstract (English)
We present a simple, model-agnostic post-processing technique for fine-grained Arabic readability classification in the BAREC 2025 Shared Task (19 ordinal levels). Our method applies conformal prediction to generate prediction sets with coverage guarantees, then computes weighted averages using softmax-renormalized probabilities over the conformal sets. This uncertainty-aware decoding improves Quadratic Weighted Kappa (QWK) by reducing high-penalty misclassifications to nearer levels. Our approach shows consistent QWK improvements of 1-3 points across different base models. In the strict track, our submission achieves QWK scores of 84.9\%(test) and 85.7\% (blind test) for sentence level, and 73.3\% for document level. For Arabic educational assessment, this enables human reviewers to focus on a handful of plausible levels, combining statistical guarantees with practical usability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。