解决多用户个性化视觉语言模型偏好趋同问题
Group Preference Collapse in Personalized Multimodal Large Language Models

- 分离用户画像与偏好表征,用原型+残差分解偏好
- 在多个模型上减少偏好坍缩,提升个体差异敏感度
- 适合需要精准个性化推荐的场景
个性化多模态大语言模型旨在生成用户特定响应,但现有方法主要依赖群体级信息,忽视用户间偏好差异。我们发现群体偏好坍缩现象:多用户个性化模型在生成过程中因偏好信号被抑制和使用不可靠,逐渐趋于主流群体选择。为此提出PrefMoE,一种以偏好为中心的框架,将稳定画像信息与偏好相关表征分离。该方法将偏好分解为共享原型与个性化残差,通过不平衡感知学习、反事实伪用户增强和残差去相关,保留个体化残差,并通过独立的LoRA适配路径分别处理画像与偏好因素。在多个MLLM骨干网络上的实验表明,PrefMoE显著提升偏好敏感性个性化,同时大幅减少偏好坍缩。项目页:https://prefmoe.github.io/
原文摘要 · Abstract (English)
Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。