arXiv:2501.18698cs.CV2025-01被引 1

对比大模型与专用模型在行人重识别中的表现,发现大模型有潜力但问题明显。

Human Re-ID Meets LVLMs: What can we expect?

  • 用提示工程和数据规范流程测试四大主流大视觉语言模型
  • 大模型在准确率上接近专用模型,但存在严重错误回答现象
  • 适合关注多模态融合与跨领域应用的研究者参考

大型视觉语言模型(LVLMs)在内容生成、虚拟助手及多模态搜索等任务中被视为突破性进展。然而,其在特定领域性能常被质疑,尤其相较于专为该领域设计的先进方法。本文在行人重识别(ReID)任务中,将ChatGPT-4o、Gemini-2.0-Flash、Claude 3.5 Sonnet和Qwen-VL-Max与专用的PersonViT模型进行对比,使用Market1501数据集。评估流程涵盖数据预处理、提示工程与指标选择,从相似度分数、分类准确率及精确率、召回率、F1值和AUC等多个维度分析结果。结果显示LVLM具备一定优势,但也暴露出频繁产生灾难性错误的问题,需进一步研究。最后建议融合传统方法与大模型技术,以实现性能显著提升。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have been regarded as a breakthrough advance in an astoundingly variety of tasks, from content generation to virtual assistants and multimodal search or retrieval. However, for many of these applications, the performance of these methods has been widely criticized, particularly when compared with state-of-the-art methods and technologies in each specific domain. In this work, we compare the performance of the leading large vision-language models in the human re-identification task, using as baseline the performance attained by state-of-the-art AI models specifically designed for this problem. We compare the results due to ChatGPT-4o, Gemini-2.0-Flash, Claude 3.5 Sonnet, and Qwen-VL-Max to a baseline ReID PersonViT model, using the well-known Market1501 dataset. Our evaluation pipeline includes the dataset curation, prompt engineering, and metric selection to assess the models' performance. Results are analyzed from many different perspectives: similarity scores, classification accuracy, and classification metrics, including precision, recall, F1 score, and area under curve (AUC). Our results confirm the strengths of LVLMs, but also their severe limitations that often lead to catastrophic answers and should be the scope of further research. As a concluding remark, we speculate about some further research that should fuse traditional and LVLMs to combine the strengths from both families of techniques and achieve solid improvements in performance.

行人重识别大模型多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。