arXiv:2606.09038cs.AI2026-06

首次系统分析个性化大模型的安全风险与应对策略

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

论文配图:Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs
图 1 · 摘自论文原文
  • 从用户表征、个性化范式到评估体系构建统一框架
  • 发现安全评估忽略用户关系,长期风险难捕捉
  • 覆盖提示、微调、强化学习等全链条风险与对策

大语言模型通过适应用户偏好、上下文和长期历史实现了日益个性化的交互,但其个性化机制也扩展了安全边界,现有研究未系统探讨这一交集。本文首次提出面向安全的个性化大模型综合综述,从用户表征、个性化范式、评估三方面组织分析,并建立统一的安全风险分类。在表征层面,分析多样化用户表示带来的风险;在主流范式中,剖析提示、检索增强、参数微调、强化学习、专家混合(MoE)、剪枝、代理框架及多模态个性化中的固有漏洞,并提出全生命周期缓解策略。此外,识别出跨范式的共性安全风险。进一步总结个性化数据集与评估方法,通过OpenClaw案例分析个性化代理生态的部署趋势。分析揭示三项结构性不足:安全评估视为用户无关而非关系型,个性化技术孤立分析而非组合评估,评估框架无法捕捉涌现的长期风险。通过联合考察个性化表征、范式、风险、防御与评估,提供发展安全个性化大模型的统一框架,并指明未来研究方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the safety landscape in ways not systematically addressed by existing literature. Existing reviews typically focus either on personalization or safety, leaving their intersection largely unexplored. We present the first comprehensive, safety-aware review of personalized LLMs. We organize personalization along three dimensions-user representation, personalization paradigm, and evaluation-and introduce a unified taxonomy of safety risks. At the representation level, we analyze risks arising from diverse user representations. Across mainstream personalization paradigms, we delineate vulnerabilities inherent to prompting, retrieval augmentation, parameter fine-tuning, reinforcement learning, Mixture-of-Experts (MoE), pruning, agent frameworks, and multimodal personalization, and synthesize mitigation strategies across the model lifecycle. Beyond these fine-grained risks, we characterize paradigm-agnostic safety risks arising from personalized adaptation. We further summarize personalized datasets and evaluation methodologies. Through a case study of OpenClaw, we analyze deployment trends in personalized agent ecosystems. Our analysis reveals three structural inadequacies in existing research: safety is evaluated as user-invariant rather than relational, personalization techniques are analyzed in isolation rather than in composition, and evaluation frameworks cannot capture emergent long-term risks. By jointly examining personalized representations, personalization paradigms, safety risks, defenses, and evaluation methods, we provide a unified framework for developing safe personalized LLMs and highlight key directions for future research.

个性化安全风险LLM综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。