通过对比学习与集成方法,提升跨语言词汇难度预测的准确性。
Improving Lexical Difficulty Prediction with Context-Aligned Contrastive Learning and Ridge Ensembling

- 用双目标对比学习对齐上下文表示,捕捉跨语言共性与差异。
- 学习到的表示能准确反映词汇难度的有序关系,提升预测稳定性。
- 集成模型有效缓解单个模型偏差,适合多语言阅读评估场景。
词汇难度预测是语言学习与可读性评估的基础问题,需估计不同母语背景下的词汇难易程度。现有方法依赖仅回归训练与标量监督,未显式构建表征空间,难以捕捉跨语言对齐与难度的序数结构。为此,我们提出上下文对齐对比回归,结合岭回归集成与两个互补目标:跨视图上下文对齐与序数软对比学习。在三个母语(L1)数据集上的实验表明:(i) 对比目标提升了跨语言表征对齐,同时保留语言特异性;(ii) 学习到的表征能有效捕获词汇难度的序数结构;(iii) 集成策略有效缓解了个体模型的系统性偏差,使各难度层级表现更稳定。
原文摘要 · Abstract (English)
Lexical difficulty prediction is a fundamental problem in language learning and readability assessment, requiring models to estimate word difficulty across different first-language (L1) backgrounds. However, existing approaches rely on regression-only training with scalar supervision, which does not explicitly structure the representation space, limiting their ability to capture cross-lingual alignment and ordinal difficulty. To mitigate these issues, we propose Context-Aligned Contrastive Regression, which integrates Ridge regression ensemble with two complementary objectives, i.e., Cross-View Context and Ordinal Soft Contrastive Learning. Experiments on three L1 datasets show that (i) contrastive objectives improve cross-lingual representation alignment while preserving language-specific nuances, (ii) the learned representations capture the ordinal structure of lexical difficulty, and (iii) the ensemble effectively mitigates systematic biases of individual models, leading to more stable performance across difficulty levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。