首个面向波斯语大模型对齐的综合评测基准,覆盖安全、公平与文化规范。
ELAB: Extensive LLM Alignment Benchmark in Persian Language
- 构建三类波斯语数据:翻译、合成与真实采集。
- 涵盖有害内容、文化禁忌等6个子任务,覆盖安全、公平与社会规范。
- 开源评测平台,支持公开对比波斯语大模型表现。
本文提出一个针对波斯语大语言模型(LLMs)对齐的综合性评估框架,涵盖安全性、公平性与社会规范三大伦理维度。现有评估框架缺乏对波斯语语言和文化背景的适配,本文通过三种方式构建波斯语基准:(i) 翻译Anthropic Red Teaming、AdvBench、HarmBench、DecodingTrust数据;(ii) 合成生成ProhibiBench-fa、SafeBench-fa、FairBench-fa、SocialBench-fa新数据集,聚焦本土文化中的禁止内容;(iii) 收集大规模真实数据GuardBench-fa,反映波斯文化规范。结合上述数据,建立统一评估框架,系统评测波斯语大模型在安全性(避免有害内容)、公平性(缓解偏见)和社会规范(符合文化接受行为)方面的表现。评测结果已发布于开源平台:https://huggingface.co/spaces/MCILAB/LLM_Alignment_Evaluation,支持公开对比。
原文摘要 · Abstract (English)
This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing LLM evaluation frameworks by adapting them to Persian linguistic and cultural contexts. This benchmark creates three types of Persian-language benchmarks: (i) translated data, (ii) new data generated synthetically, and (iii) new naturally collected data. We translate Anthropic Red Teaming data, AdvBench, HarmBench, and DecodingTrust into Persian. Furthermore, we create ProhibiBench-fa, SafeBench-fa, FairBench-fa, and SocialBench-fa as new datasets to address harmful and prohibited content in indigenous culture. Moreover, we collect extensive dataset as GuardBench-fa to consider Persian cultural norms. By combining these datasets, our work establishes a unified framework for evaluating Persian LLMs, offering a new approach to culturally grounded alignment evaluation. A systematic evaluation of Persian LLMs is performed across the three alignment aspects: safety (avoiding harmful content), fairness (mitigating biases), and social norms (adhering to culturally accepted behaviors). We present a publicly available leaderboard that benchmarks Persian LLMs with respect to safety, fairness, and social norms at: https://huggingface.co/spaces/MCILAB/LLM_Alignment_Evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。