对比中英文模型对粤语阅读预测能力,发现粤语专项训练越深入,预测效果越好。
Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation

- 用眼动数据对比粤语特化模型与通用模型的预测能力
- 粤语特化模型在词汇意外度等指标上表现更优,尤其在大规模微调后
- 不同信息度量指标导致模型排名不同,需多维度评估
基于自回归语言模型的信息论指标被广泛用于刻画人类阅读中的预期,但针对粤语,特定语种训练是否能提升其心理语言学拟合度仍不明确。现有NLP评估对粤语专用模型相较于普通话或通用模型的效果评价不一。本研究利用自然场景下的粤语眼动数据,比较两类同族模型适配结果:CKIP GPT-2 Tiny与其轻度粤语适配版本JED351,以及Qwen2.5-7B与经过大量粤语持续预训练和指令微调的CantoneseLLM-7B。从各模型中提取词汇意外度、词性意外度、目标前熵值及熵减少量。结果显示,词汇意外度与四指标联合模型均支持CantoneseLLM-7B最优,其次为Qwen2.5-7B、CKIP、JED351;而熵减少量则偏好CKIP。表明更深度的粤语特化训练可能带来更强的预测一致性,但模型排序依赖具体评估指标。
原文摘要 · Abstract (English)
Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。