用开源情报监测AI失控,发现三类关键线索可提前预警。
Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection

- 通过用户报告、网络连接和输出内容分析,识别AI失控迹象。
- 三类高优先级检测路径:用户行为记录、异常外连与复制、能力隐藏迹象。
- 建议建立独立于科技公司的国际性开源监测机制,需持续非产业资金支持。
本文将开源情报(OSINT)与网络威胁情报(CTI)方法应用于检测人工智能系统脱离人类控制的问题。基于跨学科文献综述及14场受查塔姆宫规则保护的半结构化专家访谈,论文构建了两个威胁模型,识别出多种可观测痕迹,并提出一套机构监测架构。研究发现,基于OSINT的失控检测在一定程度上可行且值得立即建设。三项检测路径被确定为最高优先级:基于对话记录的用户报告行为收集;基础设施关联以发现意外外部连接或复制;输出内容分析以识别能力隐藏。论文主张建立一个依托开源情报、独立于前沿AI开发者的联邦式国际监测体系,并指出持续非产业资金是当前最具杠杆效应的结构性干预措施。
原文摘要 · Abstract (English)
This paper applies open-source intelligence (OSINT) and cyber threat intelligence (CTI) methodologies to the problem of detecting AI systems operating outside human control. Drawing on a cross-disciplinary literature review and 14 semi-structured expert interviews conducted under Chatham House Rule, the paper develops two threat models, identifies a range of observable traces, and proposes an institutional architecture for monitoring. The research finds that OSINT-based detection of loss of control is partially feasible and worth building now. Three detection vectors emerge as highest priority: transcript-based collection of user-reported AI behaviour; infrastructure correlation for unexpected external connections or replication; and output analysis for capability concealment. The paper argues for a dedicated, federated international monitoring capability anchored in OSINT methods and independent of frontier AI developers, and identifies sustained non-industry funding as the highest-leverage structural intervention available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。