提出脑机-大模型系统路由安全审计框架,防范神经信号注入攻击。
Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents
- 设计路由安全审计契约,通过日志记录与分层验证保障决策可信
- 实测显示确认+溯源可将误接受率降至0.000,显著提升系统鲁棒性
- 适用于脑控工具代理场景,尤其关注神经信号安全性与防御机制
脑机接口到智能体的链路将解码的神经活动转化为工具使用授权通道,暴露了一种新型攻击面——脑提示注入:信号侧扰动、仅上下文注入及自适应双解码攻击均能改变路由动作,而脑电侧或文本侧监测仍无法察觉。该栈的路由安全取决于审计日志可观测性,而非解码准确率或一致性。我们定义了路由安全审计契约:最小日志模式、分母层级结构和终点规范,并证明了审计模式分离定理及C3依赖分解;干净一致性和边际鲁棒性无法识别控制C3路由的关键联合项。作为契约上的校准层,我们在非原始脑电确认通道上应用分拆共形校准,报告在明确威胁类型矩阵下的误接受率前沿。在5,400个事件的EEGMMI原生左右指令控制实验中,采用无害工具模拟、种子/案例分母,证明溯源可阻断C2路由(0.000);一致+溯源阻断C3翻转(1.000);确认+溯源完全阻断(0.000)。共形前沿在α=0.005时达到误接受率FAR 0.000、正常效用0.150;α=0.10时为FAR 0.119、效用0.452,于采集隔离下成立;攻击者可控的确认通道使边界失效至≈1。受试者聚类自助法在60名受试者上验证区间;跨架构(TinyEEGNet, EEGNetV4)与容量扫描结果表明组内饱和。
原文摘要 · Abstract (English)
BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call \emph{brain-prompt injection}: signal-side perturbations, context-only injections, and adaptive dual-decoder attacks can all change the routed action while EEG-side or text-side monitors remain blind. Route safety in this stack depends on what the audit log can observe, not on decoder accuracy or agreement alone. We define a Route-Safety Audit Contract: a minimal log schema, denominator hierarchy, and endpoint specification, and prove an audit-schema separation theorem together with a C3 attacked-dependence decomposition; clean agreement and marginal robustness do not identify the joint term that controls C3 routing. As a calibration layer on top of the contract, we apply split-conformal calibration to a non-oracle EEG confirmation channel and report the resulting false-accept frontier under an explicit threat-archetype matrix. We instantiate the contract on EEGMMI native left/right command-control over 5{,}400 events, harmless tool stubs, and seed/case denominators. Provenance blocks C2 routes ($0.000$); agreement-plus-provenance routes C3 flips ($1.000$); confirmation-plus-provenance routes them ($0.000$). The conformal frontier reaches FAR $0.000$ at clean utility $0.150$ for $α=.005$ and FAR $0.119$ at clean utility $0.452$ for $α=.10$ under acquisition isolation; an attacker-controllable confirmation channel breaks the bound to $\approx\!1$. Subject-cluster bootstrap confirms these intervals on $60$ subjects; cross-architecture (TinyEEGNet, EEGNetV4) and capacity-sweep results show within-regime saturation. Mediation and confirmation reduce risk; they are not intent certificates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。