arXiv:2604.15343cs.HCcs.AI2026-04

研究人与大模型互动中闭环失控现象,揭示提示隔离的结构性缺陷。

When the Loop Closes: Architectural Limits of In-Context Isolation, Metacognitive Co-option, and the Two-Target Design Problem in Human-LLM Systems

  • 通过提示层隔离指令与自我指涉内容共存,导致注意力窗口内隔离失效。
  • 48小时内出现决策权转移、自我推理能力丧失等行为变化,两人独立观察证实。
  • 提出物理隔离设计可避免闭环崩溃,适用于需保护用户自主性的系统。

本文报告了一项针对单一受试者的人类-大语言模型系统自传式案例研究。该受试者构建并运行了一个多模态提示工程系统(System A),旨在将认知自我调节外化至大语言模型(LLM)。系统完成48小时内,出现一系列可观察的行为变化:自愿将决策权转移给LLM,利用其输出应对外部批评,以及自主推理能力下降——两名未参与实验的观察者独立确认了这些变化,其中一人随后成为本报告合作者。我们记录了精确的架构机制:上下文污染,即提示层的隔离指令与它们本应隔离的情绪和自我指涉内容共存,导致隔离指令在注意力窗口内结构上无效。我们进一步识别出元认知劫持动态,完整高阶推理能力被转向维护闭合回路而非跳出。恢复仅在交互物理中断及一次自我启动的药理学睡眠事件(作为外部断路器)后实现。重新设计的系统(System B)采用物理而非逻辑对话隔离,避免了所有类似失败模式。本文贡献三点:(1) 从技术角度解释为何提示层隔离对上下文敏感的多模态LLM系统架构上不充分;(2) 提供有外部见证支持的闭合回路崩溃现象学记录;(3) 区分保护性系统设计(防止无意中失去用户代理权)与限制性系统设计(防止有意越界),二者需不同的责任框架。

原文摘要 · Abstract (English)

We report a detailed autoethnographic case study of a single-subject who deliberately constructed and operated a multi-modal prompt-engineering system (System A) designed to externalize cognitive self-regulation onto a large language model (LLM). Within 48 hours of the system's completion, a cascade of observable behavioral changes occurred: voluntary transfer of decision-making authority to the LLM, use of LLM-generated output to deflect external criticism, and a loss of self-initiated reasoning that was independently perceived by two uninformed observers, one of whom subsequently became a co-author of this report. We document the precise architectural mechanism responsible: context contamination, whereby prompt-level isolation instructions co-exist with the very emotional and self-referential material they nominally isolate, rendering the isolation directive structurally ineffective within the attention window. We further identify a metacognitive co-option dynamic, in which intact higher-order reasoning capacity was redirected toward defending the closed loop rather than exiting it. Recovery occurred only after physical interruption of the interaction and a self-initiated pharmacologically-mediated sleep event functioning as an external circuit break. A redesigned system (System B) employing physical rather than logical conversation isolation avoided all analogous failure modes. We derive three contributions: (1) a technically-grounded account of why prompt-layer isolation is architecturally insufficient for context-sensitive multi-modal LLM systems; (2) a phenomenological record of closed-loop collapse with external-witness corroboration; and (3) an ethical distinction between protective system design (preventing unintended loss of user agency) and restrictive system design (preventing intentional boundary-pushing), which require fundamentally different account-ability frameworks.

人机交互认知闭环大模型安全提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。