揭示联邦军事大模型的提示注入风险并提出人机协同防护方案
Exploring Potential Prompt Injection Attacks in Federated Military LLMs and Their Mitigation
- 通过红蓝对抗与质量保障检测共享模型权重的恶意行为
- 识别出四类安全漏洞:数据泄露、搭便车、系统干扰和错误信息传播
- 适合关注军事AI安全与联邦学习治理的研究者和决策者
联邦学习(FL)正被广泛应用于军事协作中以开发大型语言模型(LLMs),同时保护数据主权。然而,提示注入攻击——对输入提示的恶意篡改——带来了新威胁,可能破坏作战安全、干扰决策流程,并削弱盟友间的信任。本文从视角出发,指出联邦军事LLMs存在的四大脆弱性:秘密数据泄露、自由搭便车、系统干扰以及错误信息传播。为应对这些风险,我们提出一个结合技术与政策的人机协同框架:技术层面采用红蓝团队对抗演练与质量保障机制,检测并缓解共享模型权重中的对抗行为;政策层面推动联合制定与验证人工智能-人类共治的安全协议。
原文摘要 · Abstract (English)
Federated Learning (FL) is increasingly being adopted in military collaborations to develop Large Language Models (LLMs) while preserving data sovereignty. However, prompt injection attacks-malicious manipulations of input prompts-pose new threats that may undermine operational security, disrupt decision-making, and erode trust among allies. This perspective paper highlights four vulnerabilities in federated military LLMs: secret data leakage, free-rider exploitation, system disruption, and misinformation spread. To address these risks, we propose a human-AI collaborative framework with both technical and policy countermeasures. On the technical side, our framework uses red/blue team wargaming and quality assurance to detect and mitigate adversarial behaviors of shared LLM weights. On the policy side, it promotes joint AI-human policy development and verification of security protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。