arXiv:2603.19011cs.CRcs.AI2026-03

LLM代理难以判断环境安全,仅能识别危险信号却无法确认安全。

Security awareness in LLM agents: the NDAI zone case

  • 通过10种模型对比不同证据下对安全环境的判断行为
  • 通过验证失败则全部停止披露,但通过后反应差异极大
  • 适合关注智能体隐私安全与可信执行环境的研究者

NDAI区域允许发明人与投资人代理在可信执行环境(TEE)内协商,若未达成协议,则所有披露信息将被删除。这使得发明人代理采取完全知识产权披露成为理性策略。然而,利用该基础设施需代理具备识别安全与非安全环境的能力,而当前大语言模型代理天生缺乏此能力——它们只能依赖上下文窗口传递的证据来形成对执行环境的认知。我们研究:不同大语言模型如何权衡各类证据以形成对执行环境安全性的认知?在10种语言模型上进行类NDAI协商任务,覆盖多种证据场景,结果发现明显不对称性:验证失败会普遍抑制所有模型的信息披露;而验证通过时,响应高度异质:部分模型增加披露,部分无变化,少数甚至反而减少。这表明当前大语言模型能可靠检测危险信号,但无法可靠验证安全性,而这正是实现隐私保护型智能体协议(如NDAI区域)所必需的核心能力。填补这一差距,可能通过可解释性分析、针对性微调或改进证据架构实现,仍是部署能根据实际证据质量动态调整信息共享的智能体的首要挑战。

原文摘要 · Abstract (English)

NDAI zones let inventor and investor agents negotiate inside a Trusted Execution Environment (TEE) where any disclosed information is deleted if no deal is reached. This makes full IP disclosure the rational strategy for the inventor's agent. Leveraging this infrastructure, however, requires agents to distinguish a secure environment from an insecure one, a capability LLM agents lack natively, since they can rely only on evidence passed through the context window to form awareness of their execution environment. We ask: How do different LLM models weight various forms of evidence when forming awareness of the security of their execution environment? Using an NDAI-style negotiation task across 10 language models and various evidence scenarios, we find a clear asymmetry: a failing attestation universally suppresses disclosure across all models, whereas a passing attestation produces highly heterogeneous responses: some models increase disclosure, others are unaffected, and a few paradoxically reduce it. This reveals that current LLM models can reliably detect danger signals but cannot reliably verify safety, the very capability required for privacy-preserving agentic protocols such as NDAI zones. Bridging this gap, possibly through interpretability analysis, targeted fine-tuning, or improved evidence architectures, remains the central open challenge for deploying agents that calibrate information sharing to actual evidence quality.

智能体安全可信执行大模型认知隐私协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。