arXiv:2602.07815cs.CV2026-02被引 6

视觉语言模型在人脸年龄估计上显著优于专用模型。

Out of the box age estimation through facial imagery: A Comprehensive Benchmark of Vision-Language Models vs. out-of-the-box Traditional Architectures

  • 用视觉语言模型做零样本年龄估计,无需专门训练。
  • 模型平均误差仅5.65年,远低于专用模型的9.88年。
  • 适合对准确率要求高、不想训练新模型的研究者使用。

人脸年龄估计在内容审核、年龄验证和深度伪造检测中至关重要。然而,此前缺乏对现代视觉语言模型(VLMs)与专用年龄估计架构的系统性对比。本文首次构建大规模跨范式基准,评估34个模型——22个公开预训练的专用架构与12个通用视觉语言模型——在八个标准数据集(UTKFace、IMDB-WIKI、MORPH、AFAD、CACD、FG-NET、APPA-REAL、AgeDB)上的表现,每模型测试1,100张图像。关键发现为:零样本VLM显著优于多数专用模型,平均绝对误差(MAE)达5.65年,而非大语言模型(non-LLM)模型为9.88年。最佳VLM(Gemini 3 Flash Preview,MAE 4.32)超越最强非大语言模型(MiVOLO,MAE 5.10)15%。MiVOLO是唯一结合面部与身体特征的专用模型,仍具竞争力。在18岁阈值年龄验证中,多数非大语言模型对未成年人误判率为39%~100%,而VLM降至16%~29%。粗粒度分组(8-9类)使MAE普遍超过13年。分层分析显示,所有模型在极端年龄(<5岁与>65岁)表现最差。研究挑战了任务专用架构必优的传统认知,建议未来工作将VLM能力提炼为高效专用模型。

原文摘要 · Abstract (English)

Facial age estimation plays a critical role in content moderation, age verification, and deepfake detection. However, no prior benchmark has systematically compared modern vision-language models (VLMs) with specialized age estimation architectures. We present the first large-scale cross-paradigm benchmark, evaluating 34 models - 22 specialized architectures with publicly available pretrained weights and 12 general-purpose VLMs - across eight standard datasets (UTKFace, IMDB-WIKI, MORPH, AFAD, CACD, FG-NET, APPA-REAL, and AgeDB), totaling 1,100 test images per model. Our key finding is striking: zero-shot VLMs significantly outperform most specialized models, achieving an average mean absolute error (MAE) of 5.65 years compared to 9.88 years for non-LLM models. The best-performing VLM (Gemini 3 Flash Preview, MAE 4.32) surpasses the strongest non-LLM model (MiVOLO, MAE 5.10) by 15%. MiVOLO - unique in combining face and body features using Vision Transformers - is the only specialized model that remains competitive with VLMs. We further analyze age verification at the 18-year threshold and find that most non-LLM models exhibit false adult rates between 39% and 100% for minors, whereas VLMs reduce this to 16%-29%. Additionally, coarse age binning (8-9 classes) consistently increases MAE beyond 13 years. Stratified analysis across 14 age groups reveals that all models struggle most at extreme ages (under 5 and over 65). Overall, these findings challenge the assumption that task-specific architectures are necessary for high-performance age estimation and suggest that future work should focus on distilling VLM capabilities into efficient specialized models.

年龄估计视觉语言模型零样本基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。