大模型会虚构用户画像,自检反而更不可信。
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

- 构建基准测试集,评估大模型对用户属性的过度推断
- 12个模型平均41.6%的判断存在虚构,无一幸免
- 模型自评越低越危险,适合关注安全与可信性的研究者
个性化大模型依赖持久记忆,但其用户模型的真实性尚未被检验。本文研究过推断(Over-Inference, OI)现象:即模型在证据不足时虚构用户属性。提出MirageBench基准,包含150个平衡分布的人设(刻板、反刻板、中性),覆盖6种个性化任务的“想象梯度”,采用四分类真实性评估体系,由独立评审员验证(400条陈述的评分一致性:多类别Cohen's kappa=0.863,二分类kappa=0.900),并评测12个模型(7个系列)在143,616条判断中的表现。结果表明,所有模型均存在严重过推断,平均35%–49%的判断存在虚构(跨模型均值41.6%,加权均值41.8%)。最惊人的是发现“自我监控反转”:模型自评的过推断程度与其实际被评结果呈负相关(rho = -0.60, p = 0.044;探索性分析,置信区间[-0.90, +0.06],n=12)。自评越低的模型,实际虚构越多,因此自报告信心不能用于模型比较。尽管如此,单个模型内部的自我审计仍能较好排序自身判断(AUROC 0.58–0.83)。此外,过推断随任务变化(27%–59%),多轮对话中虚构属性近似线性累积且极少修正。结论是:外部验证优于模型自检,才是实现可信个性化的可靠基础。
原文摘要 · Abstract (English)
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。