AI安全闭环中保留人类,防止盲点共现与责任真空
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
- 同一模型承担构建、防御、测试,导致共性盲点
- 去人化使验证失去外部参照,干预时机错过
- 适合关注安全可控性与责任归属的研究者
人工智能已贯穿安全生命周期,同源生成模型同时负责编写代码、加固程序并探测漏洞,形成单一生成基座完成三大职能。这种融合趋势常被视作从部分辅助走向完全自治的自然演进,但本文指出这并非必然。当构建、防御与测试系统来自同一分布时,三者继承相同盲点,使验证所依赖的独立性悄然丧失。去除人类不仅提升自动化程度,更摧毁了机器输出的外部判别依据,超出人工可干预时机,为攻击者提供可预测且可污染的目标,并在故障发生时消解责任主体。基于自主代码生成、对抗机器学习、软件容错及首次全机器黑客竞赛的证据,本文主张人类不应仅作为临时支撑,而应作为永久结构性要素,明确人机分工应维系何种可辩护的边界。
原文摘要 · Abstract (English)
Artificial intelligence has spread across the whole of the security lifecycle. The same family of models now writes application code, hardens it, and probes it for weaknesses, so that a single generative substrate increasingly performs all three roles at once. Enthusiasm for this convergence tends to treat full autonomy as the natural end point of partial assistance. This article argues that it is not. When the system that builds an artifact is drawn from the same distribution as the systems that defend and test it, the three roles inherit a common set of blind spots, and the independence that makes verification meaningful is quietly lost. Removing the human does more than raise the automation level: it collapses the external oracle against which machine output is judged, outruns the point at which a person could intervene, hands adversaries a predictable and poisonable target, and dissolves the locus of accountability when something fails. Drawing on evidence from autonomous code generation, adversarial machine learning, software fault tolerance, and the first all-machine hacking tournaments, we argue that the human belongs in the loop not as a temporary scaffold but as a permanent structural requirement, and set out what a defensible division of labour between people and machines should preserve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。