测试大模型对巴西葡萄牙语方言的识别能力,发现存在语言偏见。
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese
- 通过提示工程分析四款大模型对巴西葡语方言的响应差异。
- 模型在不同方言间表现不一,反映出社会语言规则未被充分考虑。
- 研究为实现更公平的自然语言处理技术提供依据,适合关注语言公平者参考。
大语言模型生成内容时会再现多种偏见,包括方言偏见。本研究通过提示工程,考察了GPT 3.5、GPT-4o、Gemini和Sabi-2四款模型在区分巴西葡萄牙语变体时的表现,探究其是否考虑社会语言学规则。结果揭示了模型在方言识别中的系统性偏差,为构建更具公平性的自然语言处理技术提供了社会语言学依据。
原文摘要 · Abstract (English)
Different of biases are reproduced in LLM-generated responses, including dialectal biases. A study based on prompt engineering was carried out to uncover how LLMs discriminate varieties of Brazilian Portuguese, specifically if sociolinguistic rules are taken into account in four LLMs: GPT 3.5, GPT-4o, Gemini, and Sabi.-2. The results offer sociolinguistic contributions for an equity fluent NLP technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。