arXiv:2506.10182cs.CV2025-06CVPR被引 2

用低秩正则化微调语言编码器,实现个性化视觉-语言检索的高效适配

Improving Personalized Search with Regularized Low-Rank Parameter Updates

  • 仅调整语言编码器末层少量参数,通过低秩正则化实现个性化概念学习
  • 在DeepFashion2和ConCon-Chi数据集上个人查询检索准确率提升4%-22%
  • 支持多个人概念参数相加融合,适合需要快速定制模型的场景

个性化视觉-语言检索旨在仅凭少量样本识别新概念(如“我的狗Fido”)。该任务挑战在于需从少样本图像中学习新概念,并将个人知识与通用知识融合以在不同上下文中识别。本文提出通过正则化低秩方式微调视觉-语言双编码器的语言编码器末层少量参数,有效实现个性化检索,替代文本反转方法的同时保留通用知识。我们探索了多个已学个人概念参数的融合策略,发现参数相加有效。为评估微调后表示对通用知识的保持能力,引入基于视觉语言模型生成描述的图像检索准确率指标。所提方法在两个个性化图像检索基准(DeepFashion2和ConCon-Chi)上达到最先进性能,个人检索准确率较先前方法提升4%-22%。

原文摘要 · Abstract (English)

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the concept in different contexts. In this paper, we show how to effectively adapt the internal representation of a vision-language dual encoder model for personalized vision-language retrieval. We find that regularized low-rank adaption of a small set of parameters in the language encoder's final layer serves as a highly effective alternative to textual inversion for recognizing the personal concept while preserving general knowledge. Additionally, we explore strategies for combining parameters of multiple learned personal concepts, finding that parameter addition is effective. To evaluate how well general knowledge is preserved in a finetuned representation, we introduce a metric that measures image retrieval accuracy based on captions generated by a vision language model (VLM). Our approach achieves state-of-the-art accuracy on two benchmarks for personalized image retrieval with natural language queries - DeepFashion2 and ConCon-Chi - outperforming the prior art by 4%-22% on personal retrievals.

个性化检索低秩微调视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。