首个多模态个性化评测基准,揭示大模型在用户适配上的短板
MMPB: It's Time for Multi-Modal Personalization
- 构建10,000组图像-查询对,涵盖4类111个可个性化概念
- 23个主流模型测试显示多数难以保持对话一致性与偏好记忆
- 适合研究个性化多模态系统或评估模型用户适配能力的团队
视觉个性化在智能家居、医疗等面向用户的AI系统中至关重要,需使模型行为契合用户核心概念。然而,现有大型视觉语言模型(VLMs)在个体化适配方面仍缺乏深入探索。本文提出MMPB,首个系统性评测VLM个性化能力的基准,包含10,000组图像-查询对,覆盖人类、动物、物体、角色四类共111个可个性化概念,其中人类类别特别加入基于偏好的查询。我们将个性化划分为三类任务,分别考察模型不同关键属性。通过三阶段协议(概念注入、多轮对话、个性化查询),评估23种广泛使用的VLMs(含开源与闭源模型)。结果表明,大多数模型(包括部分闭源模型)在维持对话一致性、处理用户偏好及响应视觉线索方面表现不佳。分析揭示个性化挑战(如拒绝行为、长上下文遗忘)仍存显著改进空间。MMPB为未来实现真正个性化的多模态AI提供了重要洞见与坚实基础。
原文摘要 · Abstract (English)
Visual personalization is essential in user-facing AI systems such as smart homes and healthcare, where aligning model behavior with user-centric concepts is critical. However, recent large Vision-Language Models (VLMs), despite their broad applicability, remain underexplored in their ability to adapt to individual users. In this paper, we introduce MMPB, the first extensive benchmark for evaluating VLMs on personalization. MMPB comprises 10k image-query pairs and includes 111 personalizable concepts across four categories: humans, animals, objects, and characters, with the human category enriched with preference-grounded queries. We structure personalization into three main task types, each highlighting a different key property of VLMs. Using 23 widely used VLMs including both open- and closed-source models, we evaluate personalization performance via a three-stage protocol: concept injection, multi-turn dialogue, and personalized querying. Our findings indicate that most VLMs (including some closed-source models) struggle with personalization, particularly in maintaining consistency over dialogue, handling user preferences, and adapting to visual cues. Our analysis reveals that the challenges in VLM personalization (such as refusal behaviors and long-context forgetting) highlight substantial room for improvement. By identifying these limitations and offering a scalable benchmark, MMPB offers valuable insights and a solid foundation for future research toward truly personalized multi-modal AI. Project Page: aidaslab.github.io/MMPB
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。