arXiv:2606.16753cs.CLcs.AI2026-06中稿 · ACL

评测大模型对葡语欧版和巴版的偏见,发现多数模型偏爱巴版

P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs

论文配图:P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs
图 1 · 摘自论文原文
  • 构建多轮对话基准P3B3,评估葡语变体偏见
  • 多数模型严重偏向巴西葡语,可控性差异明显
  • 适合关注语言公平性和多语种模型评估的研究者

随着大型语言模型(LLMs)融入日常交流,捕捉地区语言差异对实现可靠且公平的语言使用至关重要。在葡萄牙语中,欧洲(pt-PT)与巴西(pt-BR)变体仍存在不均衡代表,尽管数据量上以pt-BR占优,但现有模型对葡萄牙语变体的偏好尚无深入研究。为此,我们提出P3B3——一个由专家构建、语言变体无关的对话提示基准,并配套评估框架,用于测量变体偏见与可控性。在多个模型上的实验表明,大多数LLMs对pt-BR表现出强烈偏见,且不同模型间可控性存在差异。结果凸显了在语言变体间实现更平衡的多语言表征的迫切需求。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-BR) varieties remain unevenly represented, with pt-BR dominating in data quantity, while LLM preference for Portuguese variants remains underexplored. To address this gap, we introduce P3B3, an expert-curated language variety agnostic benchmark of conversational prompts, along with an evaluation framework for measuring variety bias and controllability. Experiments on several models show that most LLMs exhibit a strong bias toward pt-BR, with variation in controllability across models. These results highlight the need for more balanced multilingual representation across language varieties.

语言模型葡语偏见评估多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。