构建丹麦语语法正确性评估新基准,提升模型测试严谨性
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
- 基于真实错误分析设计14类语法篡改函数
- 模型在新基准上表现下降,证明任务难度更高
- 适合评估语言模型对细微语法错误的敏感度
我们提出一个增强版丹麦语语法正确性评估基准。首先分析书面丹麦语中最常见的错误类型,据此设计14种系统性篡改函数,将错误引入原有正确句子中。为确保篡改准确性,采用人工与自动方法双重验证。该结果作为大语言模型语法判断任务的基准。实验表明,该基准比现有方法更广泛、更全面;通过引入更多样化的错误类型,显著提高任务难度,表现为大模型性能下降。同时,其更强的区分能力可更好辨别高性能与低性能模型。
原文摘要 · Abstract (English)
We present an enhanced benchmark for evaluating linguistic acceptability in Danish. We first analyze the most common errors found in written Danish. Based on this analysis, we introduce a set of fourteen corruption functions that generate incorrect sentences by systematically introducing errors into existing correct Danish sentences. To ensure the accuracy of these corruptions, we assess their validity using both manual and automatic methods. The results are then used as a benchmark for evaluating Large Language Models on a linguistic acceptability judgement task. Our findings demonstrate that this extension is both broader and more comprehensive than the current state of the art. By incorporating a greater variety of corruption types, our benchmark provides a more rigorous assessment of linguistic acceptability, increasing task difficulty, as evidenced by the lower performance of LLMs on our benchmark compared to existing ones. Our results also suggest that our benchmark has a higher discriminatory power which allows to better distinguish well-performing models from low-performing ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。