研究日语词汇难度感知差异,发现中文母语者因汉字词理解不同而判断难易度不同。
Difficult for Whom? A Study of Japanese Lexical Complexity
- 对比群体平均评分与个体评分,验证模型个性化潜力
- 群体模型在识别复杂词上表现接近个体模型,但预测个人难度困难
- 微调BERT仅带来微弱提升,个性化效果有限
词汇复杂度预测(LCP)与复杂词识别(CWI)通常假设难词对目标人群是共通的。尽管已有个性化方法尝试适配个体需求,我们通过部分复现标注验证了近期日语LCP数据集对目标群体的代表性。通过再次标注,发现中文母语者因对汉字词(Sino-Japanese vocabulary)理解差异,对词汇难度感知不同。为探索个性化可能性,我们比较了基于群体平均评分与个体评分训练的基准模型在个体层面的表现。结果表明,群体均值训练的模型在CWI任务中表现接近个体模型,但为个体实现良好LCP性能仍具挑战。我们还尝试微调BERT模型,但在所有设置下仅获得边际改进。
原文摘要 · Abstract (English)
The tasks of lexical complexity prediction (LCP) and complex word identification (CWI) commonly presuppose that difficult to understand words are shared by the target population. Meanwhile, personalization methods have also been proposed to adapt models to individual needs. We verify that a recent Japanese LCP dataset is representative of its target population by partially replicating the annotation. By another reannotation we show that native Chinese speakers perceive the complexity differently due to Sino-Japanese vocabulary. To explore the possibilities of personalization, we compare competitive baselines trained on the group mean ratings and individual ratings in terms of performance for an individual. We show that the model trained on a group mean performs similarly to an individual model in the CWI task, while achieving good LCP performance for an individual is difficult. We also experiment with adapting a finetuned BERT model, which results only in marginal improvements across all settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。