评测9种孟加拉方言在大模型问答中的偏差,发现方言差异越大表现越差。
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
- 用RAG管道生成4000组方言问题,结合大模型评判提升翻译保真度
- 19个大模型在方言数据上平均得分仅5.44(查塔贡方言),远低于标准语
- 提出可量化的偏差敏感度指标,适合安全关键场景的模型评估
大型语言模型在低资源语言的区域方言上常表现出性能偏差,但量化此类差异的框架仍匮乏。本文提出一个两阶段框架,评估九种孟加拉方言在大模型问答任务中的方言偏差。第一阶段,采用检索增强生成(RAG)管道将标准孟加拉语问题翻译为方言变体,并生成4,000组带黄金标签的问题集;由于传统翻译评价指标对非标准化方言无效,我们使用大模型作为评判者,经人工验证其相关性优于传统指标。第二阶段,在这些标注数据上对19个大模型进行基准测试,执行68,395次基于人类反馈的强化学习(RLAIF)评估,并通过多评委一致性与人工回退验证结果可靠性。结果显示,语言差异越大,模型表现越差:例如查塔贡方言得分仅为5.44/10,而坦盖尔方言为7.68/10。此外,模型规模增大并未一致缓解该偏差。本文贡献了经过验证的翻译质量评估方法、严谨的基准数据集及适用于高安全性应用的批判性偏差敏感度(CBS)指标。
原文摘要 · Abstract (English)
Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias in LLM question-answering across nine Bengali dialects. First, we translate and gold-label standard Bengali questions into dialectal variants adopting a retrieval-augmented generation (RAG) pipeline to prepare 4,000 question sets. Since traditional translation quality evaluation metrics fail on unstandardized dialects, we evaluate fidelity using an LLM-as-a-judge, which human correlation confirms outperforms legacy metrics. Second, we benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fallback. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous benchmark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。