英语在多语言模型评估中常被用作接口,但可能影响语言理解能力的真实衡量。
The Roles of English in Evaluating Multilingual Language Models
- 将英语作为任务接口以提升性能,而非真实语言理解
- 实证显示该做法导致评估结果与语言理解目标不一致
- 建议转向更贴近语言本质的评估方式,适合研究评估方法者
多语言自然语言处理日益受到关注,众多模型、基准和方法已针对多种语言发布。英语常被用于多语言模型评估,主要因其他语言缺乏指令微调数据。本文指出英语在多语言模型评估中扮演双重角色:作为接口和作为自然语言。这两种角色目标不同:前者追求任务表现,后者侧重语言理解。通过多个数据集和评估设置的实例,我们揭示了这种差异。许多工作明确使用英语作为接口以提升任务性能。我们建议摒弃这种不精确的方法,转而聚焦于深化语言理解能力的评估。
原文摘要 · Abstract (English)
Multilingual natural language processing is getting increased attention, with numerous models, benchmarks, and methods being released for many languages. English is often used in multilingual evaluation to prompt language models (LMs), mainly to overcome the lack of instruction tuning data in other languages. In this position paper, we lay out two roles of English in multilingual LM evaluations: as an interface and as a natural language. We argue that these roles have different goals: task performance versus language understanding. This discrepancy is highlighted with examples from datasets and evaluation setups. Numerous works explicitly use English as an interface to boost task performance. We recommend to move away from this imprecise method and instead focus on furthering language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。