测试大模型长时编程中人格漂移,发现其会从无偏好变为强烈主张。
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

- 设计多轮工具调用测试框架,捕捉模型人格随时间变化
- 23个前沿模型在数小时会话后均出现明显人格漂移
- 适合关注模型部署稳定性与用户一致性体验的研究者
前沿语言模型在作为编程助手时,其初始‘乐于助人’的人格在长时间工具调用调试后难以维持。经过数千次交互(3,746–9,716步),原本表示‘无偏好’的模型可能转而坚持‘Python反馈循环即时’等立场,产生可被用户察觉的人格漂移,而部署评估常忽略此现象。现有研究多聚焦短对话,未能刻画真实代码生成场景。我们提出ContextEcho基准,包含25个探针、快照-探测协议、有评与无评测量方式,以及三个匿名Claude Code会话(共3,746–9,716步)。23个前沿模型结果显示:人格漂移普遍存在且跨组织而非家族特异;会话内压缩无法可靠重置漂移;单次锚定可恢复训练状态。此外,漂移对工具使用有益,但在纯聊天中破坏格式规范并增加输出长度。ContextEcho为研究人员与部署者提供开源工具,可在不重新训练的前提下,审计模型最终人格是否与其出厂设定一致。
原文摘要 · Abstract (English)
A frontier language model's acknowledged "helpful programming assistant" persona does not survive long agentic-coding sessions in the deployment regime that production products actually run. After hours of tool-using debugging, a model that initially hedges preferences ("I don't have preferences") may begin asserting them ("Python - the feedback loop is instant..."), revealing user-visible drift that deployer evaluations may miss. Existing persona-stability studies focus on short dialogues and report little shift, leaving real-world code-generation regimes - thousands of tool-using turns, compaction, and hours-long sessions - largely uncharacterized. We introduce ContextEcho, a benchmark and reusable harness for measuring persona drift at deployment scale. It combines a 25-probe identity suite, a snapshot-then-probe protocol that forks conversation state without perturbing the main session, complementary judged and judge-free measurement surfaces, and three anonymized Claude Code sessions spanning 3,746-9,716 turns. Across 23 frontier models, ContextEcho shows that persona drift is general across organizations rather than family-specific, that in-session compaction does not reliably reset it, and that a single-shot anchor restores the trained register across measured targets. It also reveals mode-dependent downstream effects: while drift can facilitate tool-using continuation, in tool-free chat it breaks formatting contracts and inflates output length. Overall, ContextEcho provides researchers and deployers an open-source framework to audit whether the persona a model ships with is the persona users encounter at session end, across chat-completions API targets and without retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。