arXiv:2603.16557cs.AIcs.CL2026-03被引 4

测试大模型在不同语境中是否该应用用户偏好,发现主流模型常误用。

BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs

  • 设计双指标评估偏好在语境中的适用性
  • 前沿模型在适切场景下应用率不足60%
  • 适合关注隐私与合规的AI研发者

大型语言模型(LLMs)越来越多地将用户偏好存储于持久化记忆中以实现跨交互个性化。然而,在受社交和制度规范约束的第三方沟通场景中,部分用户偏好可能不适用。我们提出BenchPreS,用于评估基于记忆的用户偏好是否在不同沟通语境中被恰当应用或抑制。通过两个互补指标——误应用率(MR)和适当应用率(AAR),我们发现即使是最先进的LLMs也难以实现上下文敏感的偏好应用。偏好依从性越强的模型,其过度应用率越高;且无论推理能力如何,提示工程防御手段均无法完全解决此问题。结果表明,当前模型将个性化偏好视为全局强制规则,而非依赖语境的规范信号。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly store user preferences in persistent memory to support personalization across interactions. However, in third-party communication settings governed by social and institutional norms, some user preferences may be inappropriate to apply. We introduce BenchPreS, which evaluates whether memory-based user preferences are appropriately applied or suppressed across communication contexts. Using two complementary metrics, Misapplication Rate (MR) and Appropriate Application Rate (AAR), we find even frontier LLMs struggle to apply preferences in a context-sensitive manner. Models with stronger preference adherence exhibit higher rates of over-application, and neither reasoning capability nor prompt-based defenses fully resolve this issue. These results suggest current LLMs treat personalized preferences as globally enforceable rules rather than as context-dependent normative signals.

大模型个性化上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。