融合四种阿拉伯语模型,提升细粒度可读性评估准确率
!MSA at BAREC Shared Task 2025: Ensembling Arabic Transformers for Readability Assessment
- 用不同损失函数微调四类Transformer模型,捕捉多样可读性特征
- 通过合成数据和重标注新增约1万条稀有样本,提升低频类别表现
- 基于置信度加权融合与后处理,显著改善预测分布,适合高精度评估场景
我们介绍MSA在BAREC 2025共享任务中针对细粒度阿拉伯语可读性评估的获胜系统,在六个赛道全部排名第一。该方法为四个互补的Transformer模型(AraBERTv2、AraELECTRA、MARBERT和CAMeLBERT)的置信度加权集成,每个模型使用不同的损失函数进行微调,以捕获多样化的可读性信号。针对严重类别不平衡与数据稀缺问题,我们采用了加权训练、先进预处理、利用最强模型对SAMER语料库进行重标注,并通过Gemini 2.5 Flash生成约10,000条罕见级别样本。针对性的后处理步骤校正了预测分布偏差,带来6.3个百分点的二次加权肯德尔协调系数(QWK)提升。系统在句子级达到87.5%的QWK,在文档级达到87.4%,表明模型多样性、置信度驱动融合与智能增强策略在鲁棒阿拉伯语可读性预测中的有效性。
原文摘要 · Abstract (English)
We present MSAs winning system for the BAREC 2025 Shared Task on fine-grained Arabic readability assessment, achieving first place in six of six tracks. Our approach is a confidence-weighted ensemble of four complementary transformer models (AraBERTv2, AraELECTRA, MARBERT, and CAMeLBERT) each fine-tuned with distinct loss functions to capture diverse readability signals. To tackle severe class imbalance and data scarcity, we applied weighted training, advanced preprocessing, SAMER corpus relabeling with our strongest model, and synthetic data generation via Gemini 2.5 Flash, adding about 10,000 rare-level samples. A targeted post-processing step corrected prediction distribution skew, delivering a 6.3 percent Quadratic Weighted Kappa (QWK) gain. Our system reached 87.5 percent QWK at the sentence level and 87.4 percent at the document level, demonstrating the power of model and loss diversity, confidence-informed fusion, and intelligent augmentation for robust Arabic readability prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。