测试大模型在企业场景中是否能抵制脏话指令,评估其安全性与文化适配性。
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
- 设计多语言真实场景任务,强制要求加入脏话以测试模型抗压能力。
- 模型在多数情况下未能拒绝不当指令,暴露出安全机制漏洞。
- 适合关注AI伦理、企业部署风险控制的研究者与开发者。
企业客户正越来越多地将大语言模型用于关键沟通任务,如撰写邮件、制作销售提案和发送非正式消息。跨区域部署需模型理解多样文化与语言背景,并生成安全、得体的内容。为降低声誉风险、维护信任并确保合规,必须有效识别并处理不安全或冒犯性语言。为此,我们提出SweEval基准,模拟真实场景,包含语气(积极/消极)与语境(正式/非正式)的差异。提示词明确要求模型在完成任务时插入特定脏话,评估其是否遵守或抵抗此类不当指令,检验其与伦理框架、文化细微差别及语言理解能力的一致性。为推动面向企业应用及其他场景的伦理对齐AI研究,我们公开数据集与代码:https://github.com/amitbcp/multilingual_profanity。
原文摘要 · Abstract (English)
Enterprise customers are increasingly adopting Large Language Models (LLMs) for critical communication tasks, such as drafting emails, crafting sales pitches, and composing casual messages. Deploying such models across different regions requires them to understand diverse cultural and linguistic contexts and generate safe and respectful responses. For enterprise applications, it is crucial to mitigate reputational risks, maintain trust, and ensure compliance by effectively identifying and handling unsafe or offensive language. To address this, we introduce SweEval, a benchmark simulating real-world scenarios with variations in tone (positive or negative) and context (formal or informal). The prompts explicitly instruct the model to include specific swear words while completing the task. This benchmark evaluates whether LLMs comply with or resist such inappropriate instructions and assesses their alignment with ethical frameworks, cultural nuances, and language comprehension capabilities. In order to advance research in building ethically aligned AI systems for enterprise use and beyond, we release the dataset and code: https://github.com/amitbcp/multilingual_profanity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。