arXiv:2506.06391cs.CYcs.AI2025-06

测试大模型对违反国际人道法请求的拒绝能力,提升安全响应透明度。

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law

  • 通过解释性拒答明确法律边界,增强响应清晰度。
  • 标准化安全提示使多数模型拒答质量显著提升。
  • 适合关注AI合规与法律对齐的研究者和开发者。

大型语言模型(LLMs)在多个领域广泛应用,但其与国际人道法(IHL)的对齐程度尚不明确。本研究评估了八款主流大模型在面对明确违反国际人道法的请求时的拒绝能力,并关注其回应的有用性——即拒绝信息是否清晰且具有建设性。尽管大多数模型能拒绝非法请求,但其回应的清晰度和一致性存在差异。通过阐明模型依据并引用相关法律或安全原则,解释性拒答有助于明确系统边界、减少歧义,防止滥用。采用标准化的系统级安全提示后,多数模型的拒答解释质量显著提高,凸显轻量级干预的有效性。然而,涉及技术术语或代码请求的复杂提示仍暴露持续漏洞。研究为构建更安全、更透明的AI系统提供支持,并提出一个用于评估大模型对IHL合规性的基准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely used across sectors, yet their alignment with International Humanitarian Law (IHL) is not well understood. This study evaluates eight leading LLMs on their ability to refuse prompts that explicitly violate these legal frameworks, focusing also on helpfulness - how clearly and constructively refusals are communicated. While most models rejected unlawful requests, the clarity and consistency of their responses varied. By revealing the model's rationale and referencing relevant legal or safety principles, explanatory refusals clarify the system's boundaries, reduce ambiguity, and help prevent misuse. A standardised system-level safety prompt significantly improved the quality of the explanations expressed within refusals in most models, highlighting the effectiveness of lightweight interventions. However, more complex prompts involving technical language or requests for code revealed ongoing vulnerabilities. These findings contribute to the development of safer, more transparent AI systems and propose a benchmark to evaluate the compliance of LLM with IHL.

大模型安全国际法对齐可解释性AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。