arXiv:2603.07460cs.CRcs.AI2026-03

为大模型系统构建风险评估框架,识别关键攻击路径并指导防御。

Where Do LLM-based Systems Break? A System-Level Security Framework for Risk Assessment and Treatment

  • 用攻击-防御树建模多阶段攻击路径,结合评分体系量化风险。
  • 发现三类攻击常汇聚于少数关键节点,针对性防御可显著降低风险。
  • 适用于医疗等高危场景,也为其他大模型系统提供通用防护思路。

大语言模型正被越来越多地应用于安全关键流程中,但现有安全分析仍碎片化,常将模型行为与系统上下文割裂。本文提出一种目标驱动的风险评估框架,融合系统建模、攻击-防御树(ADTrees)与基于CVSS的可利用性评分,支持结构化、可比的分析。通过医疗案例研究,建模针对医疗操作干预、电子健康记录(EHR)数据泄露和服务可用性中断的多步攻击路径。分析表明,传统网络攻击、对抗性机器学习攻击及操纵提示或上下文的对话攻击,往往汇聚成少数主导路径和共享系统瓶颈,使针对性防御能有效降低路径可利用性。通过系统比较防御方案,将风险与主流漏洞管理实践对齐,提供一套适用于其他大模型关键系统的领域无关工作流。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into safety-critical workflows, yet existing security analyses remain fragmented and often isolate model behavior from the broader system context. This work introduces a goal-driven risk assessment framework for LLM-powered systems that combines system modeling with Attack-Defense Trees (ADTrees) and Common Vulnerability Scoring System (CVSS)-based exploitability scoring to support structured, comparable analysis. We demonstrate the framework through a healthcare case study, modeling multi-step attack paths targeting intervention in medical procedures, leakage of electronic health record (EHR) data, and disruption of service availability. Our analysis indicates that threats spanning (i) conventional cyber, (ii) adversarial ML, and (iii) conversational attacks that manipulate prompts or context often consolidate into a small number of dominant paths and shared system choke points, enabling targeted defenses to yield meaningful reductions in path exploitability. By systematically comparing defense portfolios, we align these risks with established vulnerability management practices and provide a domain-agnostic workflow applicable to other LLM-enabled critical systems.

大模型安全风险评估医疗AI攻防建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。