让系统通过对话逐步获取线索,提升找人准确率
LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification
- 基于视觉与文本上下文生成针对性提问,引导用户提供关键信息
- 在多个数据集上显著优于传统方法,对话式检索效果更优
- 适合需要动态交互的安防、刑侦等实际场景
传统基于文本的人体重识别假设目击者能一次性提供完整描述,但在真实场景中描述往往不完整或模糊。为此,我们提出新任务——交互式人体重识别(Inter-ReID),即通过对话逐步完善初始描述。为此,我们构建了一个包含多类型问题的对话数据集,通过分解个体细粒度属性来实现。我们进一步提出LLaVA-ReID模型,该模型基于视觉和文本上下文生成有针对性的问题以获取目标人物的额外信息。训练时采用前瞻策略,优先使用最具有信息量的问题作为监督信号。在Inter-ReID及传统文本重识别基准上的实验表明,LLaVA-ReID显著优于基线方法。
原文摘要 · Abstract (English)
Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often partial or vague. To address this limitation, we introduce a new task called interactive person re-identification (Inter-ReID). Inter-ReID is a dialogue-based retrieval task that iteratively refines initial descriptions through ongoing interactions with the witnesses. To facilitate the study of this new task, we construct a dialogue dataset that incorporates multiple types of questions by decomposing fine-grained attributes of individuals. We further propose LLaVA-ReID, a question model that generates targeted questions based on visual and textual contexts to elicit additional details about the target person. Leveraging a looking-forward strategy, we prioritize the most informative questions as supervision during training. Experimental results on both Inter-ReID and text-based ReID benchmarks demonstrate that LLaVA-ReID significantly outperforms baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。