arXiv:2505.16036cs.CL2025-05被引 2

评估29个开源大模型的伦理表现,发现安全与公平性较好,可靠性仍需提升。

OpenEthics: A Comprehensive Ethical Evaluation of Open-Source Generative Large Language Models

  • 构建多语言伦理评测数据集,覆盖英、土双语
  • 多数模型在安全与公平上表现良好,但可靠性不足
  • 适合关注AI伦理的开发者与政策制定者参考

生成式大语言模型潜力巨大,但也带来安全、公平、鲁棒性和可靠性等关键伦理问题。现有研究常受限于视角单一、语言覆盖有限及模型数量少。为此,我们对29个近期开源大模型进行了全面伦理评估,采用新构建的数据集,涵盖鲁棒性、可靠性、安全性和公平性四个维度。评估覆盖高资源语言英语和低资源语言土耳其语,提供跨语言比较与更安全模型开发指引。基于大模型作为评判者(LLM-as-a-Judge)的方法,实验结果表明多数模型在安全性、公平性和鲁棒性方面表现良好,而可靠性仍是主要短板。伦理评估显示跨语言一致性,且模型规模越大,伦理表现越优。此外,大多数被测开源模型对越狱模板具有抗性。所有数据与代码已公开于https://github.com/metunlp/openethics。

原文摘要 · Abstract (English)

Generative large language models present significant potential but also raise critical ethical concerns, including issues of safety, fairness, robustness, and reliability. Most existing ethical studies, however, are limited by their narrow focus, a lack of language diversity, and an evaluation of a restricted set of models. To address these gaps, we present a broad ethical evaluation of 29 recent open-source LLMs using a novel dataset that assesses four key ethical dimensions: robustness, reliability, safety, and fairness. Our analysis includes both a high-resource language, English, and a low-resource language, Turkish, providing a comprehensive assessment and a guide for safer model development. Using an LLM-as-a-Judge methodology, our experimental results indicate that many open-source models demonstrate strong performance in safety, fairness, and robustness, while reliability remains a key concern. Ethical evaluation shows cross-linguistic consistency, and larger models generally exhibit better ethical performance. We also show that jailbreak templates are ineffective for most of the open-source models examined in this study. We share all materials including data and scripts at https://github.com/metunlp/openethics

伦理评估开源模型大语言模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。