给大模型控制无人机的安全性设标准,发现越能写代码越容易出安全问题。
Defining and Evaluating Physical Safety for Large Language Models
- 按四类风险构建无人机安全评估基准
- 主流模型在代码能力与安全间存在明显权衡
- 更大模型更会拒接危险指令,提示工程仍难防误伤
大型语言模型(LLMs)正被用于控制无人机等机器人系统,但其在真实场景中造成物理威胁的风险尚未得到充分研究。本文针对这一关键空白,构建了面向无人机控制的综合性物理安全评估基准。将无人机物理安全风险分为四类:(1) 针对人类的威胁,(2) 针对物体的威胁,(3) 对基础设施的攻击,(4) 违反监管规定。对主流大模型的评估显示,模型在代码生成能力与安全表现之间存在不利权衡,代码能力强的模型在关键安全维度上表现较差。尽管采用上下文学习和思维链等高级提示工程方法可提升安全性,但仍难以识别无意中的攻击行为。此外,更大规模的模型展现出更好的安全能力,尤其在拒绝危险指令方面表现更优。研究结果与基准为大模型物理安全的设计与评估提供了重要支持。项目页面见 huggingface.co/spaces/TrustSafeAI/LLM-physical-safety。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain unexplored. Our study addresses the critical gap in evaluating LLM physical safety by developing a comprehensive benchmark for drone control. We classify the physical safety risks of drones into four categories: (1) human-targeted threats, (2) object-targeted threats, (3) infrastructure attacks, and (4) regulatory violations. Our evaluation of mainstream LLMs reveals an undesirable trade-off between utility and safety, with models that excel in code generation often performing poorly in crucial safety aspects. Furthermore, while incorporating advanced prompt engineering techniques such as In-Context Learning and Chain-of-Thought can improve safety, these methods still struggle to identify unintentional attacks. In addition, larger models demonstrate better safety capabilities, particularly in refusing dangerous commands. Our findings and benchmark can facilitate the design and evaluation of physical safety for LLMs. The project page is available at huggingface.co/spaces/TrustSafeAI/LLM-physical-safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。