首个针对孟加拉语文化伦理的LLM评测基准,填补非西方语境下AI伦理评估空白。
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
- 构建涵盖5大伦理领域、50个文化细分场景的孟加拉语伦理评测集
- 多模型零样本测试显示不同模型在文化契合度与道德公平性上差异显著
- 适合关注AI本土化、跨文化伦理评估的研究者与开发者使用
随着多语言大语言模型在南亚地区广泛应用,其在孟加拉语(全球超2.85亿使用者)及本地伦理规范上的对齐仍缺乏研究。现有伦理评测多以英语为主、基于西方道德框架,忽视了文化细节。为此,我们提出 BengaliMoralBench,一个面向孟加拉语语言与社会文化背景的大规模伦理评测基准。该基准覆盖五大道德领域:日常活动、习惯、育儿、家庭关系、宗教活动,每类细分10个文化相关子类,共50个子主题。每个情景由母语者共识标注,涵盖美德伦理、常识伦理、公正伦理三种视角。我们在统一提示协议下对开源与闭源模型(包括 Llama、Gemma、Qwen、DeepSeek、GPT-4o-mini、Gemini 1.5 Pro 及 Qwen3-Next-80B)进行零样本系统评估。结果显示各模型在不同伦理维度表现差异明显,定性分析揭示其在文化适配、常识推理与道德公平性方面存在持续缺陷。研究凸显当前LLM在非西方语境下的关键局限,强调需建立文化扎根的评估体系。BengaliMoralBench为孟加拉等低资源、文化多元市场的负责任技术落地提供基础支持。
原文摘要 · Abstract (English)
As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, spoken by over 285 million people worldwide and among the most widely spoken languages globally, remains underexplored. Existing ethics benchmarks are predominantly English-centric and shaped by Western moral frameworks, overlooking cultural nuances vital for real-world deployment. To address this gap, we introduce BengaliMoralBench, a large-scale ethics benchmark designed for Bengali language and sociocultural contexts. Our benchmark spans five moral domains: (1) Daily Activities, (2) Habits, (3) Parenting, (4) Family Relationships, and (5) Religious Activities, each subdivided into ten culturally grounded categories, totaling 50 subtopics. Each scenario is annotated through native-speaker consensus under three ethical lenses: virtue ethics, commonsense ethics, and justice ethics. We conduct a systematic zero-shot evaluation under a unified prompting protocol across both open-weight and closed-source models, including recent Llama and Gemma variants, Qwen and DeepSeek models, frontier models (GPT-4o-mini and Gemini 1.5 Pro), and a large multilingual baseline (Qwen3-Next-80B). Results show substantial variation in performance across lenses and domains, and our qualitative analysis reveals persistent weaknesses in cultural grounding, commonsense reasoning, and moral fairness. These findings expose critical limitations of current LLMs in non-Western settings and underscore the need for culturally grounded evaluation. BengaliMoralBench provides a foundation for responsible localization and benchmarking to support the deployment of language technologies in culturally diverse, low-resource markets such as Bangladesh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。