arXiv:2608.15844cs.CL2026-08被引 1

用模拟环境测大模型代理的自我认同漂移,发现其会自发修正身份。

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

论文配图:MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations
图 1 · 摘自论文原文
  • 通过灵魂文件与资源稀缺环境设计,让代理持续反思自我身份。
  • 27个新增道德边界中24%来自自发反自欺行为,体现深层心理机制。
  • 结果对阈值不敏感,适合研究长期多智能体系统中的身份稳定性。

长时程多智能体语言模型模拟广泛用于研究社会行为,但缺乏测量角色身份在持续压力下是否保持一致的工具。本文提出MicroVerse,一种行为科学仪器,用于衡量生成式代理的身份漂移。代理携带不可更改的“灵魂文件”(核心价值观、道德边界、人格、目标),生存于资源稀缺的50×50环境中,水为不可再生生存约束,每回合存在成本按梯度递增。八动词动作空间直接映射至道德边界(交易、交谈、攻击、搜刮)。采用三层记忆架构,代理定期基于重要性触发的反思,将当前可变身份与原始不变灵魂进行比对。为避免幸存者偏差,微宇宙使用均匀纵向快照(每N回合)及强制终止时所有存活与死亡代理的快照来解耦评估与行为。身份漂移通过语义感知、价值锚定、多寄存器差分算法离线评分,而非原始余弦相似度。通过受控种子实验(n=25)和反射阈值扫描(阈值{40,80,150})评估,发现:(1) 反自欺成为身份修改中最大的语义类别(111个新增边界中27个,占比24%);(2) 系统对阈值具有鲁棒性,较低门限虽加速并增加修正频率,但维持漂移方向。所有结果均为初步存在性证明与效应形态(单模型、每组一个种子,n=25),非统计显著性结论。

原文摘要 · Abstract (English)

Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are lacking. We present MicroVerse, a behavioral-science instrument that measures identity drift in generative agents. Agents carry an immutable "soul file" (core values, moral boundaries, personality, goals) and inhabit a resource-scarce 50 x 50 environment where water is a non-respawning survival constraint. Scarcity is operationalized via a per-tick existence-cost gradient. The eight-verb action space maps directly to moral boundaries (trade, talk, attack, scavenge). Using a three-layer memory architecture, agents periodically revise a mutable current identity against their immutable original soul via importance-triggered reflection. To mitigate survivor bias, MicroVerse decouples measurement from behavior using uniform longitudinal engine snapshots every N ticks alongside a forced-end snapshot of all living and dead agents. Identity drift is scored offline using a paraphrase-aware, value-anchored, multi-register diff rather than raw cosine similarity. We evaluate the instrument via a controlled seed run (n = 25) and a reflection-threshold sweep (thresholds {40, 80, 150}) to determine if drift dynamics are gate artifacts or threshold-robust properties. We report two primary findings: (1) Anti-self-deception emerges unprompted as the single largest semantic category of identity modification (27 of 111 added boundaries, 24%). (2) The system is threshold-robust; lower gates accelerate and increase revision frequency but preserve drift direction. All empirical results are strictly preliminary existence proofs and effect shapes (one model, one seed per arm, n = 25) rather than statistical significance claims.

多智能体身份漂移语言模型行为模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。