用知识图谱构建可解释的防御评估体系,提升大模型抗攻击能力
Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation
- 基于本体建模攻击与防御方法,形成结构化论证框架
- 生成人类可读、机器可处理的可信性证明文档
- 适用于工程师、数据科学家等多方角色的安全部署与审计
尽管大型语言模型(LLMs)展现出强大适应性,但在安全性、透明性和可解释性方面仍面临挑战。由于其易受对抗攻击影响,需结合对抗训练与防护机制持续保障鲁棒性。然而,管理隐含且异构的知识以持续验证鲁棒性极为困难。本文提出一种基于形式化论证的新型方法,利用本体对前沿攻击与防御手段进行形式化建模,构建可由人类阅读、机器解析的可信性证据链。通过英语语言和代码翻译任务中的实例验证该方法的有效性,并为工程实践与理论发展提供启示,服务于工程师、数据科学家、用户及审计人员。
原文摘要 · Abstract (English)
Despite the impressive adaptability of large language models (LLMs), challenges remain in ensuring their security, transparency, and interpretability. Given their susceptibility to adversarial attacks, LLMs need to be defended with an evolving combination of adversarial training and guardrails. However, managing the implicit and heterogeneous knowledge for continuously assuring robustness is difficult. We introduce a novel approach for assurance of the adversarial robustness of LLMs based on formal argumentation. Using ontologies for formalization, we structure state-of-the-art attacks and defenses, facilitating the creation of a human-readable assurance case, and a machine-readable representation. We demonstrate its application with examples in English language and code translation tasks, and provide implications for theory and practice, by targeting engineers, data scientists, users, and auditors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。