BLiSS 1.0用真实学习者语料测试模型对语言错误的自然容忍度
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
- 构建280万条真实学习者句子,生成13.7万组三元组对比
- 发现模型对自然错误更容忍,此能力与语法正确性独立
- 适合评估小模型在语言习得规律上的拟合程度
为弥合性能导向基准与认知启发模型评估之间的差距,我们提出BLiSS 1.0——学习者跨语言句法结构基准。该基准引入选择性容忍新范式,测试模型是否认为同一句子中的自然学习者错误比匹配的人工错误更合理。基准基于超过280万条自然学习者语料,提供136,867组受控三元组(修正句、学习者句、人工错误句)。在多种模型上的实验表明,选择性容忍是独立于标准语法正确性的能力,性能聚类明显由训练范式决定。这验证了BLiSS作为衡量不同训练目标如何影响模型与人类语言习得系统性模式对齐程度的稳健工具。
原文摘要 · Abstract (English)
To bridge the gap between performance-oriented benchmarks and the evaluation of cognitively inspired models, we introduce BLiSS 1.0, a Benchmark of Learner Interlingual Syntactic Structure. Our benchmark operationalizes a new paradigm of selective tolerance, testing whether a model finds a naturalistic learner error more plausible than a matched, artificial error within the same sentence. Constructed from over 2.8 million naturalistic learner sentences, BLiSS provides 136,867 controlled triplets (corrected, learner, artificial) for this purpose. Experiments on a diverse suite of models demonstrate that selective tolerance is a distinct capability from standard grammaticality, with performance clustering strongly by training paradigm. This validates BLiSS as a robust tool for measuring how different training objectives impact a model's alignment with the systematic patterns of human language acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。