arXiv:2602.05877cs.AIcs.HC2026-02被引 4

为车载大模型助手设计安全威胁分类体系,聚焦人类安全风险。

Agent2Agent Threats in Safety-Critical LLM Assistants: A Human-Centric Taxonomy

  • 基于人类受害建模构建安全资产分类体系
  • 区分恶意数据传播与触发动作两类攻击路径
  • 提供开源工具实现多阶段威胁自动发现

将基于大语言模型(LLM)的对话代理集成到车辆中,带来了智能代理、汽车安全与代理间通信交汇处的新安全挑战。这些智能助手通过Google的Agent-to-Agent(A2A)等协议与外部服务协同,形成攻击面,恶意自然语言载荷可导致驾驶员分心甚至未经授权的车辆控制。现有AI安全框架虽具基础性,但缺乏安全关键系统工程中的严格‘职责分离’标准,将保护对象(资产)与攻击路径混同。本文提出名为AgentHeLLM(LLM助手的代理危害探索)的威胁建模框架,形式化分离资产识别与攻击路径分析。引入基于伤害导向的‘受害者建模’的人类中心资产分类体系,受《世界人权宣言》启发;并建立形式化图模型,区分毒化路径(恶意数据传播)与触发路径(激活行为)。通过开源攻击路径生成工具AgentHeLLM Attack Path Generator,展示框架实用性,该工具采用双层搜索策略实现多阶段威胁自动化发现。

原文摘要 · Abstract (English)

The integration of Large Language Model (LLM)-based conversational agents into vehicles creates novel security challenges at the intersection of agentic AI, automotive safety, and inter-agent communication. As these intelligent assistants coordinate with external services via protocols such as Google's Agent-to-Agent (A2A), they establish attack surfaces where manipulations can propagate through natural language payloads, potentially causing severe consequences ranging from driver distraction to unauthorized vehicle control. Existing AI security frameworks, while foundational, lack the rigorous "separation of concerns" standard in safety-critical systems engineering by co-mingling the concepts of what is being protected (assets) with how it is attacked (attack paths). This paper addresses this methodological gap by proposing a threat modeling framework called AgentHeLLM (Agent Hazard Exploration for LLM Assistants) that formally separates asset identification from attack path analysis. We introduce a human-centric asset taxonomy derived from harm-oriented "victim modeling" and inspired by the Universal Declaration of Human Rights, and a formal graph-based model that distinguishes poison paths (malicious data propagation) from trigger paths (activation actions). We demonstrate the framework's practical applicability through an open-source attack path suggestion tool AgentHeLLM Attack Path Generator that automates multi-stage threat discovery using a bi-level search strategy.

AI安全车载AI威胁建模人机安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。