让AI识人对话,实现个性化视觉交互
Personalized Visual Instruction Tuning
- 自动生成含人物特性的对话数据,训练模型识人
- 在新基准上性能显著提升,实现精准个性化回复
- 适合家庭机器人、手机助手等需认人的场景
多模态大语言模型虽有显著进展,但仍存在‘人脸失认’问题:能泛泛聊天,却无法针对特定个体进行个性化对话。这限制了其在手机视觉助手或家庭机器人等个性化场景的应用。本文提出个性化视觉指令微调(PVIT),通过一套自动化数据生成管道,利用视觉专家、图像生成模型和多模态大模型,构建包含个性化对话的训练数据。该方法使模型能识别图像中特定人物并开展连贯对话。为评估个性化能力,我们构建了涵盖多种题型与难度的基准P-Bench。实验表明,使用该数据集微调后,模型的个性化表现显著提升。
原文摘要 · Abstract (English)
Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness". Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeting at specific individuals. This deficiency hinders the application of MLLMs in personalized settings, such as tailored visual assistants on mobile devices, or domestic robots that need to recognize members of the family. In this paper, we introduce Personalized Visual Instruction Tuning (PVIT), a novel data curation and training framework designed to enable MLLMs to identify target individuals within an image and engage in personalized and coherent dialogues. Our approach involves the development of a sophisticated pipeline that autonomously generates training data containing personalized conversations. This pipeline leverages the capabilities of various visual experts, image generation models, and (multi-modal) large language models. To evaluate the personalized potential of MLLMs, we present a benchmark called P-Bench, which encompasses various question types with different levels of difficulty. The experiments demonstrate a substantial personalized performance enhancement after fine-tuning with our curated dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。