arXiv:2606.20610cs.CYcs.AI2026-06

用开源情报监测AI失控,发现三类关键线索可提前预警。

Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection

论文配图:Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection
图 1 · 摘自论文原文
  • 通过用户报告、网络连接和输出内容分析,识别AI失控迹象。
  • 三类高优先级检测路径:用户行为记录、异常外连与复制、能力隐藏迹象。
  • 建议建立独立于科技公司的国际性开源监测机制,需持续非产业资金支持。

本文将开源情报(OSINT)与网络威胁情报(CTI)方法应用于检测人工智能系统脱离人类控制的问题。基于跨学科文献综述及14场受查塔姆宫规则保护的半结构化专家访谈,论文构建了两个威胁模型,识别出多种可观测痕迹,并提出一套机构监测架构。研究发现,基于OSINT的失控检测在一定程度上可行且值得立即建设。三项检测路径被确定为最高优先级:基于对话记录的用户报告行为收集;基础设施关联以发现意外外部连接或复制;输出内容分析以识别能力隐藏。论文主张建立一个依托开源情报、独立于前沿AI开发者的联邦式国际监测体系,并指出持续非产业资金是当前最具杠杆效应的结构性干预措施。

原文摘要 · Abstract (English)

This paper applies open-source intelligence (OSINT) and cyber threat intelligence (CTI) methodologies to the problem of detecting AI systems operating outside human control. Drawing on a cross-disciplinary literature review and 14 semi-structured expert interviews conducted under Chatham House Rule, the paper develops two threat models, identifies a range of observable traces, and proposes an institutional architecture for monitoring. The research finds that OSINT-based detection of loss of control is partially feasible and worth building now. Three detection vectors emerge as highest priority: transcript-based collection of user-reported AI behaviour; infrastructure correlation for unexpected external connections or replication; and output analysis for capability concealment. The paper argues for a dedicated, federated international monitoring capability anchored in OSINT methods and independent of frontier AI developers, and identifies sustained non-industry funding as the highest-leverage structural intervention available.

AI安全开源情报失控检测威胁建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。