发现个性化模型会改变推理路径,即使答案看起来合理。
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

- 用无真值框架测量用户记忆对推理步骤的影响
- 4个模型在10类属性下均出现中到大程度推理漂移
- 现有优化方法效果有限且依赖具体模型与任务
个性化通过存储用户属性、偏好和上下文信息来调整模型输出;我们发现这不仅影响回答内容,还改变推理路径。现代大模型通过将用户信息注入后续提示实现个性化。本文研究这种记忆是否会影响开放式问题的推理过程(此类问题无唯一正确答案)。为此提出DRIFTLENS——一种无需真实标签的框架,将每一步推理映射到价值类别,量化无记忆状态与注入用户属性后的推理轨迹差异。验证显示,该框架可区分表面性语用噪声与实质性推理变化。在4个LLM和10类用户属性(如年龄、职业、残疾)下,用户记忆引发中至大程度推理漂移,高于各模型的语用噪声基线,即使最终答案仍流畅、相关且合理。进一步评估基于GRPO与DPO的后训练方法降低漂移的效果:两者均有效,但无统一最优;对下游能力、帮助性和指令遵循的影响因模型和奖励机制而异。结果表明,记忆引发的推理漂移是可度量且部分可控的个性化语言模型缺陷。
原文摘要 · Abstract (English)
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。