用多模态模型统一解决纹身、素描等刑侦图像检索难题
Bridging the Modality Gap in Forensic Image Retrieval

- 用大语言模型自动生成图文描述,实现跨模态匹配
- 融合文本与图像特征后,检索准确率显著提升
- 适合刑侦人员处理素描、口述描述等复杂场景
自动化图像检索在现代法医分析中日益关键,支持依赖高效视觉证据比对的调查流程。以往研究主要集中于多模态检索系统的开发与优化,但对这些技术在多样化真实场景中的法医适用性评估不足。本文提出一个统一的检索框架,适配四项关键法医任务:(1) 以纹身图像为查询进行纹身检索;(2) 根据人工专家提供的文本描述检索纹身,模拟目击者口头描述的常见情形;(3) 从手绘素描中检索纹身;(4) 从法医人脸素描中检索人脸。系统利用多模态大语言模型(MLLM)自动为所有查询和图库图像生成结构化文本描述,随后采用句向量嵌入进行文本比对。我们通过纯视觉嵌入、纯文本嵌入以及融合文本与图像相似度得分的多模态融合策略进行评估,该策略结合了各任务相关的先进视觉特征提取器。多模态融合在视觉信息有限或嘈杂的场景下(如素描、部分纹身或碎片化证词)持续提升检索精度与鲁棒性。本研究凸显了统一多模态检索流水线的法医价值,并展示了现代MLLM如何将传统依赖人工专家分析的挑战性任务实现自动化。结果表明,多模态检索是支持涉及纹身、面部合成及证人描述等调查工作的有力工具。
原文摘要 · Abstract (English)
Automated image retrieval plays an increasingly critical role in modern forensic analysis, supporting investigative workflows that rely on efficient comparison of visual evidence. While prior work has focused primarily on developing and optimizing multimodal retrieval systems, limited attention has been paid to evaluating the forensic applicability of these technologies across diverse real-world scenarios. In this study, we present a unified retrieval framework adapted to four key forensic tasks: (1) tattoo image retrieval given a tattoo query image; (2) tattoo retrieval guided by human-expert textual descriptions, modelling the common situation where a witness verbally describes a tattoo; (3) tattoo retrieval from hand-drawn sketches; and (4) face retrieval from forensic face sketches. Our system leverages a multimodal large language model (MLLM) to automatically generate structured textual descriptions for all queries and gallery images, followed by sentence-transformer embedding for text-based comparison. We evaluate retrieval using visual-only embeddings, text-only embeddings and a multimodal fusion strategy that combines text- and image-based similarity scores derived from state-of-the-art visual feature extractors relevant to each task. The fusion of modalities consistently improves retrieval precision and robustness, especially in scenarios where visual information is limited or noisy (e.g., sketches, partial tattoos, or fragmented witness statements). This work highlights the forensic value of a unified multimodal retrieval pipeline and demonstrates how modern MLLMs can operationalize challenging forensic tasks that traditionally rely on manual expert analysis. Our results position multimodal retrieval as a promising tool for supporting investigative workflows involving tattoos, facial composites, and witness descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。