首个针对四指掌纹验证的多模态大模型评测基准,揭示提示词设计对识别效果的关键影响。
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

- 构建首个四指掌纹多模态模型评测集SLAPBench,基于NIST SD302b数据集
- 提示词类型决定模型表现:任务描述导致多数模型误判率近100%,相似度评分提升性能
- 发现模型能力差异显著,部分模型出现近似随机判断或反向性能,需警惕偏见风险
四指SLAP掌纹是单手食指、中指、无名指和小指的平面活体扫描指纹,用于边境管控与执法身份验证。目前尚无基准评估多模态大语言模型(MLLMs)从SLAP图像中进行身份验证的能力。本文提出SLAPBench,首个基于NIST SD302b数据集的MLLM四指掌纹验证基准,包含7,832对样本(176对同源,7,656对非同源)。评估了四个开源模型(InternVL3-8B、Qwen2.5-VL-7B、Qwen3-VL-8B、Gemma-3-12B)和一个专有模型Claude Opus 4.8,采用零样本、任务描述和相似度评分三类提示。提示词显著影响验证行为:任务描述提示下,所有开源模型误接受率接近100%;仅Claude Opus 4.8在二分类提示下保持稳定,错误接受率20.2%最优。相似度评分可避免崩溃现象,暴露模型能力差距:Claude AUC达0.953,Gemma-3-12B为0.837,InternVL3-8B倒置至0.590,Qwen2.5-VL-7B接近随机(0.567)。而Qwen3-VL-8B达到完美分离(AUC=1.000),经控制实验确认非分辨率捷径所致,但可能受近似重复样本干扰。公平性分析显示,歧视随弱化趋势加剧。本研究建立首个专用基准,揭示提示设计主导模型崩溃,模型能力决定偏差程度。
原文摘要 · Abstract (English)
Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs (176 mated, 7,656 non-mated). We evaluate four open-source MLLMs (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, Gemma-3-12B) and the proprietary Claude Opus 4.8 under zero-shot, task-description, and similarity-scoring prompts. Prompting governs verification behavior. Task-description prompting collapses all four open-source models to near-100% False Accept Rate (FAR), and Gemma-3-12B collapses under zero-shot as well; Claude Opus 4.8 alone resists collapse under both binary prompts, giving the best binary result (FAR = 20.2%). Similarity scoring removes collapse across the open-source models and exposes wide capability gaps: Claude reaches AUC = 0.953 and Gemma-3-12B 0.837, while InternVL3-8B is inverted (AUC = 0.590) and Qwen2.5-VL-7B near random (0.567). Qwen3-VL-8B attains perfect separation (AUC = 1.000), which we treat as a diagnostic rather than as capability: SD302b holds one SLAP capture per finger position, so mated pairs are cross-resolution. A matched-resolution control leaves the perfect score intact, ruling out the resolution shortcut; what cannot be excluded within SD302b is near-duplicate detection, since a mated pair is one capture rendered twice. A fairness probe over gender, race, and age suggests disparity grows as discrimination weakens. SLAPBench establishes the first SLAP-specific MLLM baseline and shows that prompting governs collapse while model capability governs discrimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。