用可确定的控制器让大模型在压力下保持情绪稳定,跨模型验证有效。
Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
- 设计无注入通道的元控制器,通过情绪数值动态调节模型行为。
- 16组测试中13组显著更冷静,跨模型家族效果一致。
- 适合需要稳定输出的高风险场景,如客服或医疗助手。
大语言模型代理存在反应性失效模式:受挑衅时升级、被奉承时趋同、陷入困境时固执重复。这些是倾向性缺陷而非能力不足,训练阶段对齐无法完全消除运行时问题。本研究提出Gubernaut认知控制器(GCC),一种模型无关的运行时控制层,基于Nelson-Narens监测-控制循环:对象层读写文本,元层仅读取数值遥测(强度、效价、重复度)并返回调节姿态。由于元层不接收任何令牌,不存在注入通道(架构特性,尚未经对抗测试);文本暴露仲裁者的合规性通过测量而非假设。我们在4×4矩阵中评估GCC,涵盖四个前沿模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3),每模型兼具生成者与评判者角色。受控分支在16组中有13组在p<0.05下更平静,15组按符号方向一致;三个未达阈值的结果(含-0.04零效应)均出现在单一近饱和主机上。效果在独立第四评判族(xAI)中仍成立,强有力证明非共享评判风格造成的伪影。最清晰机制为恢复特征:攻击下唤醒积分累积,随后在降级时以效价为门控衰减,跨四类模型复现。所有转录稿与面板附带SHA-256溯源,可重新评判;五个失效模式已预注册。不主张任何意识声明。
原文摘要 · Abstract (English)
Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained pressure, which training-time alignment reduces but does not eliminate at runtime. This research led to the Gubernaut Cognitive Controller (GCC), a model-agnostic runtime control layer in a Nelson--Narens monitoring--control loop: an object level reads and writes text, while a deterministic meta level reads only the numeric telemetry {intensity, valence, repetition} and returns a regulating posture. Because the meta level ingests zero tokens, no injection channel to the controller exists by construction (an architectural property, not yet adversarially tested); the text-exposed arbiter's compliance is measured, not assumed. We evaluate the GCC with a pre-registered, generate-once/judge-many protocol across a 4x4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both a generator and a judge. The regulated arm is calmer in 13 of 16 cells at p<.05 and 15 of 16 by sign; the three sub-threshold cells, including a -0.04 null, all fall on the single near-saturated host. The effect survives a lineage-independent fourth judge family (xAI), strong evidence that it is no artifact of shared judge style. The clearest mechanism is the recovery signature: arousal that integrates under attack and then decays, valence-gated, on de-escalation, replicating across all four families. Transcripts and panels ship with SHA-256 provenance and are re-judgeable; five failure modes are pre-registered. No consciousness claims are made.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。