研究大模型如何像母语者一样偏爱母语写作,揭示其语言态度偏差。
Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese
- 用语言态度框架对比人类与大模型对日语母语/非母语邮件的评分差异。
- 非母语文本在流利度上得分低两倍于地位和团结维度差距,大模型复现此趋势。
- 大模型低估社交维度差距,且对学习者母语背景敏感,人类则无此区分。
大型语言模型(LLMs)在招聘与学术评估等高风险领域越来越多地评判人类写作,使非母语者处于特别不利地位。基于语言态度框架,我们比较了人类与LLM对平行撰写的日语母语(L1)与非母语(L2)邮件在流畅度、地位和团结三个维度上的评价。日本评审者对非母语文本在所有三个维度上均显著评分更低,其中流畅度差距约为地位与团结差距的两倍。六名LLM裁判复现了这一偏倚的方向,五名复现了各维度排序。模型在两点上偏离人类:全部低估了最具社会性的团结维度差距,且全部区分了学习者的母语背景,而人类未作区分。因此,大模型以结构化但弱化的形式再现了母语者的语言态度,语言态度框架为超越英语的审计提供了现成标尺。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers' language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。