用上千人描述风格训练模型,让AI生成更多样的人像文本描述。
Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
- 通过聚类人类描述风格,为每种风格设计提示词。
- 在风格空间均匀采样,生成更丰富的描述原型。
- 大幅提升文本到图像重识别的泛化能力,适合需要多样化数据的应用。
文本到图像人像重识别(Text-to-Image Person ReID)旨在根据文本描述检索目标人物的图像。该任务的主要挑战在于大规模数据库的人工标注成本高,影响模型泛化能力。现有方法利用多模态大语言模型(MLLM)自动生成行人图像描述,但生成的描述风格单一。为此,本文提出人类标注者建模(HAM)方法,使MLLM能够模仿数千名人类标注者的描述风格。具体而言,首先从人类文本描述中提取风格特征并进行聚类,将相似风格的描述归入同一簇;随后为每个簇设计提示词,并采用提示学习模拟不同标注者的风格。此外,构建风格特征空间,通过均匀采样获得更具代表性的聚类原型,进一步提升描述多样性。最终,利用HAM自动标注大规模文本到图像重识别数据库。大量实验表明,该数据库显著提升了ReID模型的泛化性能。
原文摘要 · Abstract (English)
Text-to-image person re-identification (ReID) aims to retrieve the images of an interested person based on textual descriptions. One main challenge for this task is the high cost in manually annotating large-scale databases, which affects the generalization ability of ReID models. Recent works handle this problem by leveraging Multi-modal Large Language Models (MLLMs) to describe pedestrian images automatically. However, the captions produced by MLLMs lack diversity in description styles. To address this issue, we propose a Human Annotator Modeling (HAM) approach to enable MLLMs to mimic the description styles of thousands of human annotators. Specifically, we first extract style features from human textual descriptions and perform clustering on them. This allows us to group textual descriptions with similar styles into the same cluster. Then, we employ a prompt to represent each of these clusters and apply prompt learning to mimic the description styles of different human annotators. Furthermore, we define a style feature space and perform uniform sampling in this space to obtain more diverse clustering prototypes, which further enriches the diversity of the MLLM-generated captions. Finally, we adopt HAM to automatically annotate a massive-scale database for text-to-image ReID. Extensive experiments on this database demonstrate that it significantly improves the generalization ability of ReID models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。