arXiv:2509.12662cs.CL2025-09EMNLP被引 50

用对话生成和交互优化文本搜人,免人工标注还能更准。

Chat-Driven Text Generation and Interaction for Person Retrieval

  • 通过多轮对话模拟生成伪标签,无需人工写描述。
  • 推理时动态优化模糊查询,提升真实场景适应性。
  • 全流程免标注,适合实际监控系统部署。

基于文本的人像搜索(TBPS)通过自然语言描述从大规模数据库中检索图像,在监控应用中具有重要意义。然而,高质量文本标注过程耗时费力,限制了其可扩展性和实际应用。为此,我们提出两个互补模块:多轮文本生成(MTG)与多轮文本交互(MTI)。MTG 利用多模态大模型模拟对话,自动生成细粒度、多样化的视觉描述伪标签,实现无监督生成。MTI 在推理阶段通过动态对话式推理,优化用户查询,有效处理模糊、不完整或歧义的描述,这正是真实场景中的常见问题。二者结合形成统一的免标注框架,显著提升检索精度、鲁棒性与可用性。大量实验表明,该方法在无需人工标注的前提下达到竞争性或更优性能,为 TBPS 系统的规模化与实用化铺平道路。

原文摘要 · Abstract (English)

Text-based person search (TBPS) enables the retrieval of person images from large-scale databases using natural language descriptions, offering critical value in surveillance applications. However, a major challenge lies in the labor-intensive process of obtaining high-quality textual annotations, which limits scalability and practical deployment. To address this, we introduce two complementary modules: Multi-Turn Text Generation (MTG) and Multi-Turn Text Interaction (MTI). MTG generates rich pseudo-labels through simulated dialogues with MLLMs, producing fine-grained and diverse visual descriptions without manual supervision. MTI refines user queries at inference time through dynamic, dialogue-based reasoning, enabling the system to interpret and resolve vague, incomplete, or ambiguous descriptions - characteristics often seen in real-world search scenarios. Together, MTG and MTI form a unified and annotation-free framework that significantly improves retrieval accuracy, robustness, and usability. Extensive evaluations demonstrate that our method achieves competitive or superior results while eliminating the need for manual captions, paving the way for scalable and practical deployment of TBPS systems.

文本搜人对话生成免标注多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。