arXiv:2410.05306cs.CRcs.AI2024-10中稿 · the AI Act Worksho…被引 5

构建可解释框架,助大模型合规并抵御对抗攻击

Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs

  • 用本体+保障论证+事实表构建可解释合规框架
  • 支持工程师验证模型在对抗攻击下的鲁棒性
  • 适合监管方与开发者用于AI安全合规评估

大语言模型易被滥用且易受安全威胁,引发重大安全与合规风险。欧盟《人工智能法案》虽要求特定场景下确保AI鲁棒性,但因缺乏标准、模型复杂及新兴漏洞而难以落地。本文提出一种基于本体、保障论证与事实表的框架,帮助工程师与利益相关方理解并记录大模型在对抗鲁棒性方面的合规性与安全性。该方法旨在确保大模型符合监管要求,并具备应对潜在威胁的能力。

原文摘要 · Abstract (English)

Large language models are prone to misuse and vulnerable to security threats, raising significant safety and security concerns. The European Union's Artificial Intelligence Act seeks to enforce AI robustness in certain contexts, but faces implementation challenges due to the lack of standards, complexity of LLMs and emerging security vulnerabilities. Our research introduces a framework using ontologies, assurance cases, and factsheets to support engineers and stakeholders in understanding and documenting AI system compliance and security regarding adversarial robustness. This approach aims to ensure that LLMs adhere to regulatory standards and are equipped to counter potential threats.

大模型安全合规框架对抗鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。