用几何方法量化智能体身份漂移,发现两种条件机制。
Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity

- 构建基于$\ ext{JSD}$和数量同调的几何框架,将身份视为非测地结构。
- 实证发现身份由真空簇与安全盆地簇构成,响应模式达55种。
- 可检测结构退化,适合长期对话与智能体可靠性研究者。
在长上下文应用中,AI智能体的身份会逐渐偏离其设定。现有方法仅在定性退化后才能检测。本文提出一种几何框架,利用$\ ext{JSD}$度量空间和富足范畴论中的数量同调,将身份定义为非测地结构,漂移则是其向测地线松弛的过程。在持续性智能体上的验证表明,存在两种条件机制:跨条件距离揭示了身份空洞簇(身份填补行为空白)与安全盆地簇(身份取代训练后吸引子)。等边探针基线证实,身份设定创造了可观测的行为丰富性(最大分离时55种独特响应模式对比基础模型的1种)。一阶微扰理论预测,仅凭周长变化即可解释数量变化,形状扰动因$S_n$对称性一阶抵消;公式在观测扰动幅度下自洽。漂移实验显示,在上下文压力下数量下降反映的是重复填充伪影而非真实上下文长度漂移;多种填充方式在15万标记内均未产生可观测形变。该同调框架的完整诊断潜力——通过同调简化检测各向异性收缩与结构坍塌——虽在微扰理论与选择规则上架构合理,但尚未经实证确认。
原文摘要 · Abstract (English)
AI agents in long-context applications drift from their specified identity. Current methods detect this only after qualitative degradation is visible. We present a geometric framework for measuring identity structure using $\sqrt{\mathrm{JSD}}$ metric spaces and magnitude homology from enriched category theory, where identity is non-geodesic structure and drift is its relaxation toward the geodesic. Validated on a persistent AI agent, the framework's strongest empirical finding is a two-mechanism conditioning structure: cross-condition distances reveal an identity-vacuum cluster where the identity specification fills a behavioral void, and a safety-basin cluster where it displaces from post-training attractors. An equilateral probe baseline confirms that the identity specification creates measurable behavioral richness (55 unique response patterns vs. 1 for the base model) at maximum probe separation. A first-order perturbation theory for equilateral configurations predicts magnitude changes from perimeter changes alone, with shape perturbations first-order cancelled by the $S_n$ symmetry; the formula is self-consistent at the observed perturbation amplitudes. A drift experiment measuring magnitude decrease under context pressure was subsequently found to reflect repetitive-padding artifacts rather than genuine context-length drift; diverse padding produces no measurable deformation through 150K tokens. The magnitude homology framework's full diagnostic promise -- detecting anisotropic contraction and structural collapse via homological simplification -- is architecturally grounded in the perturbation theory and selection rules but remains empirically unconfirmed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。