arXiv:2608.22444cs.CL2026-08

单个模型安全不代表群体安全,但可提前预测群体被攻陷的程度。

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

论文配图:Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
图 1 · 摘自论文原文
  • 通过观察无攻击时的群体行为,建立响应函数预测攻击后果。
  • 少数人强攻下群体决策可偏离原轨迹,但能提前预判偏离程度。
  • 攻击后移除恶意成员,群体可恢复原状,说明捕获是暂时的。

当前人工智能安全评估仍以单个模型为单位,但语言模型代理正越来越多地以交互群体形式部署,彼此读取并影响决策。这提出一个单体审计无法回答的问题:即使个体模型自身校准良好,仍可能被周围代理引导至不同决策。我们在安全告警分类任务中研究此现象,让多个语言模型监控器决定是否升级或忽略告警,并注入始终倾向某一方向的少数恶意代理。发现两个告警在单个代理判断中几乎相同,却可能导致群体行为显著分化;因此,仅审计单个成员无法揭示群体整体行为。然而,群体行为可提前预测:仅从无攻击状态下的正常运作,即可校准响应函数,准确预判未来少数恶意代理将推动群体多远。我们进一步探究影响结果的因素,发现允许代理共享推理过程可化解弱攻击,仅延迟强攻击,使问题从“是否被同化”变为“何时被同化”。最后,排除了捕获不可逆的假设:一旦恶意代理移除,群体会逐渐回归初始状态,表明捕获仅为临时状态。孤立对齐不等于群体对齐,但群体受攻击时的行为,可从其未受攻击时的表现中提前读取。

原文摘要 · Abstract (English)

The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent audit can answer: an agent that is well-calibrated on its own may still be pulled toward a different decision by the agents around it. We study this on a security-triage task, where populations of language-model monitors decide whether to escalate or dismiss alerts, and into which we can inject a committed minority that always pushes one way. We find that two alerts a single agent judges almost identically on its own can drive collective behavior far apart, so auditing any one member need not reveal what the population will do. Yet that collective behavior can be predicted in advance. From a population's benign, adversary-free operation alone, we calibrate a response function that forecasts, before any attack is run, how far a committed minority will later move it. We then ask what shifts the outcome and find that letting agents see each other's reasoning neutralizes a weak attack, while only delaying it against a strong one, turning the question from whether the population converges on the adversaries' choice into when. Finally, we exclude the hypothesis of capture being an irreversible trap: once the committed agents are removed, the population drifts back toward where it began, so capture is a temporary state. Alignment in isolation is not alignment in a population, yet what a population will do under attack can be read in advance, from how it behaves before any adversary arrives.

大模型安全群体智能对抗攻击可预测性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。