arXiv:2607.28818cs.AIcs.CL2026-07被引 1

测试AI伙伴长期互动中角色崩溃与行为漂移问题,发现现有模型难以维持一致性。

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

论文配图:Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
图 1 · 摘自论文原文
  • 设计ANCHOR审计框架,分离评估角色扮演与对话轨迹记忆能力。
  • 平均轨迹准确率仅44.4%,用户状态回忆接近随机水平。
  • 结果受模型、角色特征和评估者影响,需区分评估维度。

随着AI伴侣越来越多地参与重复社交互动,用户依赖稳定的角色设定与共同历史,但局部合理的回复并不保证其持续性。本文研究两种可观测的长期失效现象:'角色崩溃'(即部署角色、边界、价值观或风格的丧失)与'行为漂移'(这些属性的渐进或反复退化)。我们提出ANCHOR,一种可控的合成审计方法,分别测量角色执行与对话轨迹回忆。实验包含2,008次对话,覆盖27个角色、9种交互频率、3种生成记忆设置及4个评估模型。身份探测器结合密封的102题问卷与逐轮判断,轨迹探测器则对35个对话库中的110道校准反事实问题进行评分。结果显示,所有评估模型与配置均无法可靠维持任一维度:轨迹准确率平均仅为44.4%,用户状态回忆接近四选一随机水平;未发现任何上下文条件或记忆设置能稳定缓解这些问题。问卷保留率因模型和角色维度而异,且与逐轮行为不一致,对评估者选择敏感。表明当前系统尚无法可靠支持长期陪伴连续性,审计必须区分角色执行、轨迹回忆、评估者来源与部署环境,而非合并为单一信任或稳定性指标。

原文摘要 · Abstract (English)

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.

AI伴侣角色一致性行为漂移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。