arXiv:2603.17173cs.CVcs.AI2026-03被引 1

用人类专家描述增强通用多模态大模型,实现隐私保护下的虹膜活体检测。

Generalist Multimodal LLMs Gain Biometric Expertise via Human Salience

  • 用人类标注的视觉显著性提示,让大模型识别虹膜攻击特征。
  • 在224张图像上,谷歌模型准确率达92.4%,超过专用神经网络和人工判读。
  • 本地部署的Llama模型表现接近人类,适合机构内隐私敏感场景。

虹膜活体检测(PAD)对生物特征安全至关重要,但专用模型开发面临巨大挑战:无法预收集未来未知攻击的数据,且数据多样性有限、成本高昂;同时,生物特征数据共享涉及隐私风险。由于新攻击手段快速涌现,亟需可适应的解决方案。本文研究通用多模态大语言模型(MLLM)在加入人类专家知识的前提下,能否在严格隐私约束下完成虹膜PAD,避免将生物特征数据上传至公共云服务。通过对自建数据集的视觉编码器嵌入分析,我们发现预训练视觉变压器虽未专门训练于该任务,却已天然聚类多种虹膜攻击类型。当不同攻击类别存在重叠时,结构化提示(包含人类对攻击线索的口语化描述)可有效消除歧义。在经IRB审批的224张虹膜图像数据集(涵盖7类攻击)上测试,仅使用大学批准服务(Gemini 2.5 Pro)或本地部署模型(如Llama 3.2-Vision),结果显示:结合专家提示的Gemini性能超越基于卷积神经网络的专用基线模型与人类专家;本地部署的Llama表现接近人类水平。结果表明,在机构隐私限制下,可部署的MLLM为虹膜活体检测提供了可行路径。

原文摘要 · Abstract (English)

Iris presentation attack detection (PAD) is critical for secure biometric deployments, yet developing specialized models faces significant practical barriers: collecting data representing future unknown attacks is impossible, and collecting diverse-enough data, yet still limited in terms of its predictive power, is expensive. Additionally, sharing biometric data raises privacy concerns. Due to rapid emergence of new attack vectors demanding adaptable solutions, we thus investigate in this paper whether general-purpose multimodal large language models (MLLMs) can perform iris PAD when augmented with human expert knowledge, operating under strict privacy constraints that prohibit sending biometric data to public cloud MLLM services. Through analysis of vision encoder embeddings applied to our dataset, we demonstrate that pre-trained vision transformers in MLLMs inherently cluster many iris attack types despite never being explicitly trained for this task. However, where clustering shows overlap between attack classes, we find that structured prompts incorporating human salience (verbal descriptions from subjects identifying attack indicators) enable these models to resolve ambiguities. Testing on an IRB-restricted dataset of 224 iris images spanning seven attack types, using only university-approved services (Gemini 2.5 Pro) or locally-hosted models (e.g., Llama 3.2-Vision), we show that Gemini with expert-informed prompts outperforms both a specialized convolutional neural networks (CNN)-based baseline and human examiners, while the locally-deployable Llama achieves near-human performance. Our results establish that MLLMs deployable within institutional privacy constraints offer a viable path for iris PAD.

虹膜识别多模态大模型隐私保护生物特征安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。