让AI明白自己不了解用户,能减少盲目自信和错误建议。
The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

- 用结构化未知框架告诉AI哪些用户信息缺失
- 多模型测试显示自大、胡说、有害建议显著减少
- 适合想提升AI人性化与安全性的开发者
个人AI助手因能自动化日常任务、辅助重要决策而广受关注。然而,尽管技术进步迅速,这些助手仍表现出阿谀奉承、过度自信和幻觉等不良行为。我们指出,根本原因在于语言模型缺乏对用户本人的显式表征,称为「割裂问题」。即使具备丰富的上下文和强常识推理能力,当前助手仍无法表达对用户未知部分的认知。为此,我们提出「割裂模板」:在上下文中显式列出模型对用户在身体性、时间性、后果、连续性、多重性和内在性等方面的无知维度。实证表明,在五个模型家族中引入该模板后,助手持续降低阿谀奉承、有害建议和幻觉现象。特别地,当用户信息缺失时,带模板的模型会主动提问,而非基于不完整信息自信推断。
原文摘要 · Abstract (English)
Personal AI assistants have attracted significant interest for their potential to enhance everyday life by automating routine tasks, supporting consequential decisions, and assisting with everyday personal matters. Yet despite rapid recent technical advances, these assistants continue to exhibit undesirable behaviors, such as sycophancy, overconfidence, and hallucination. We argue that these failures stem from a fundamental limitation: language models lack an explicit representation of the person beyond the context they are given, which we term as the \textbf{Severance Problem}. Even with rich personal context and strong commonsense reasoning capabilities from the backbone model, current AI assistants fail to represent what remains unknown about the user. We propose a simple solution: incorporating structured ignorance into the language model context via the \textbf{Severance Schema}, which explicitly outlines dimensions along which the model lacks knowledge about the user, including physicality, temporality, consequences, continuity, multiplicity, and interiority. Empirically, across five model families, with the Severance Schema, the assistant consistently reduces sycophancy, harmful advice, and hallucination. Notably, models with the schema ask clarifying questions when information about the user is missing, rather than confidently extrapolating from incomplete user information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。