arXiv:2603.26015cs.CVcs.AI2026-03

零样本评估大模型识别人脸年龄,发现通用模型表现堪比专业模型。

VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation

  • 不微调直接用大视觉语言模型做零样本年龄估计。
  • 在UTKFace和FG-NET上达到与专业模型相当的准确率。
  • 适合关注公平性、零样本推理与跨领域应用的研究者。

从人脸图像中进行人类年龄估计是计算机视觉中的挑战性任务,在生物识别、医疗健康和人机交互中有重要应用。传统深度学习方法需大量标注数据和领域特定训练,而大视觉语言模型(LVLMs)提供了零样本估计的可能性。本研究对GPT-4o、Claude 3.5 Sonnet和LLaMA 3.2 Vision在两个基准数据集UTKFace和FG-NET上进行了无微调的零样本评估。采用8项指标(包括MAE、MSE、RMSE、MAPE、MBE、$R^2$、CCC和±5年准确率)验证,结果显示通用型LVLM在零样本设置下可实现具有竞争力的性能。研究揭示了图像质量与人口子群体间的性能差异,强调了多模态推理中公平性的重要性。该工作构建了可复现的基准,推动LVLM在法医学、健康监测与人机交互等现实场景中的应用,同时指出了提示敏感性、可解释性、计算成本与公平性等仍需解决的问题。

原文摘要 · Abstract (English)

Human age estimation from facial images represents a challenging computer vision task with significant applications in biometrics, healthcare, and human-computer interaction. While traditional deep learning approaches require extensive labeled datasets and domain-specific training, recent advances in large vision-language models (LVLMs) offer the potential for zero-shot age estimation. This study presents a comprehensive zero-shot evaluation of state-of-the-art Large Vision-Language Models (LVLMs) for facial age estimation, a task traditionally dominated by domain-specific convolutional networks and supervised learning. We assess the performance of GPT-4o, Claude 3.5 Sonnet, and LLaMA 3.2 Vision on two benchmark datasets, UTKFace and FG-NET, without any fine-tuning or task-specific adaptation. Using eight evaluation metrics, including MAE, MSE, RMSE, MAPE, MBE, $R^2$, CCC, and $\pm$5-year accuracy, we demonstrate that general-purpose LVLMs can deliver competitive performance in zero-shot settings. Our findings highlight the emergent capabilities of LVLMs for accurate biometric age estimation and position these models as promising tools for real-world applications. Additionally, we highlight performance disparities linked to image quality and demographic subgroups, underscoring the need for fairness-aware multimodal inference. This work introduces a reproducible benchmark and positions LVLMs as promising tools for real-world applications in forensic science, healthcare monitoring, and human-computer interaction. The benchmark focuses on strict zero-shot inference without fine-tuning and highlights remaining challenges related to prompt sensitivity, interpretability, computational cost, and demographic fairness.

年龄估计零样本大模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。