首个针对指纹分析的多模态大模型评测基准,验证了模型在指纹细节识别上的潜力。
FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- 构建涵盖7个数据集的FPBench评测体系,覆盖8类指纹任务。
- 零样本与思维链提示下,开源模型性能提升7%-39%。
- 适合生物特征识别、法医学及多模态模型研究者参考。
多模态大语言模型(MLLMs)具备复杂数据分析、视觉问答、生成与推理能力,但其在生物特征数据方面的应用仍不充分。本文系统评估了MLLMs对指纹图像中精细结构与纹理特征的理解能力。为此,我们设计了首个综合性评测基准FPBench,用于评估20个开源与专有模型在7个真实与合成数据集上的表现,涵盖8项生物特征与法医学任务(如纹路分析、指纹验证、真伪分类等),采用零样本与思维链提示策略。同时,我们在部分开源模型上微调视觉与语言编码器,验证领域适应效果。结果表明,微调可使性能提升7%-39%。该工作为指纹领域的基础模型发展迈出第一步。代码已公开于https://github.com/Ektagavas/FPBench。
原文摘要 · Abstract (English)
Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their ability to analyze biometric data is relatively underexplored. In this work, we investigate the effectiveness of MLLMs in understanding fine structural and textural details present in fingerprint images. To this end, we design a comprehensive benchmark, FPBench, to evaluate 20 MLLMs (open-source and proprietary models) across 7 real and synthetic datasets on a suite of 8 biometric and forensic tasks (e.g., pattern analysis, fingerprint verification, real versus synthetic classification, etc.) using zero-shot and chain-of-thought prompting strategies. We further fine-tune vision and language encoders on a subset of open-source MLLMs to demonstrate domain adaptation. FPBench is a novel benchmark designed as a first step towards developing foundation models in fingerprints. Our findings indicate fine-tuning of vision and language encoders improves the performance by 7%-39%. Our codes are available at https://github.com/Ektagavas/FPBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。