arXiv:2503.10192cs.SEcs.CL2025-03被引 8

测试三款AI模型在西语和巴斯克语中的安全与偏见问题

Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

  • 人工红队测试670轮对话,评估模型安全与偏见
  • 阿尔亚萨拉曼德拉模型问题率最高达50.6%
  • 揭示多语言AI系统仍存重大安全风险

AI领导权争夺战中,美国OpenAI与中国的DeepSeek是主要竞争者。为应对这一趋势,西班牙政府提出ALIA计划,构建公开透明的公共人工智能基础设施,包含小型语言模型以支持西班牙语及共官方语言巴斯克语。本文报告了红队测试结果,10名专家对OpenAI o3-mini、DeepSeek R1和ALIA Salamandra三款最新模型进行手动测试,聚焦偏见与安全问题。基于670轮对话分析,所有模型均存在漏洞,出现偏见或不安全回复的比例从o3-mini的29.5%到Salamandra的50.6%不等。研究凸显了开发可靠可信AI系统的持续挑战,尤其针对西班牙语与巴斯克语的应用场景。

原文摘要 · Abstract (English)

The battle for AI leadership is on, with OpenAI in the United States and DeepSeek in China as key contenders. In response to these global trends, the Spanish government has proposed ALIA, a public and transparent AI infrastructure incorporating small language models designed to support Spanish and co-official languages such as Basque. This paper presents the results of Red Teaming sessions, where ten participants applied their expertise and creativity to manually test three of the latest models from these initiatives$\unicode{x2013}$OpenAI o3-mini, DeepSeek R1, and ALIA Salamandra$\unicode{x2013}$focusing on biases and safety concerns. The results, based on 670 conversations, revealed vulnerabilities in all the models under test, with biased or unsafe responses ranging from 29.5% in o3-mini to 50.6% in Salamandra. These findings underscore the persistent challenges in developing reliable and trustworthy AI systems, particularly those intended to support Spanish and Basque languages.

红队测试多语言AI安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。