零样本评估大模型识别人脸年龄,发现通用模型表现堪比专业模型。
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
- 不微调直接用大视觉语言模型做零样本年龄估计。
- 在UTKFace和FG-NET上达到与专业模型相当的准确率。
- 适合关注公平性、零样本推理与跨领域应用的研究者。
从人脸图像中进行人类年龄估计是计算机视觉中的挑战性任务,在生物识别、医疗健康和人机交互中有重要应用。传统深度学习方法需大量标注数据和领域特定训练,而大视觉语言模型(LVLMs)提供了零样本估计的可能性。本研究对GPT-4o、Claude 3.5 Sonnet和LLaMA 3.2 Vision在两个基准数据集UTKFace和FG-NET上进行了无微调的零样本评估。采用8项指标(包括MAE、MSE、RMSE、MAPE、MBE、$R^2$、CCC和±5年准确率)验证,结果显示通用型LVLM在零样本设置下可实现具有竞争力的性能。研究揭示了图像质量与人口子群体间的性能差异,强调了多模态推理中公平性的重要性。该工作构建了可复现的基准,推动LVLM在法医学、健康监测与人机交互等现实场景中的应用,同时指出了提示敏感性、可解释性、计算成本与公平性等仍需解决的问题。
原文摘要 · Abstract (English)
Human age estimation from facial images represents a challenging computer vision task with significant applications in biometrics, healthcare, and human-computer interaction. While traditional deep learning approaches require extensive labeled datasets and domain-specific training, recent advances in large vision-language models (LVLMs) offer the potential for zero-shot age estimation. This study presents a comprehensive zero-shot evaluation of state-of-the-art Large Vision-Language Models (LVLMs) for facial age estimation, a task traditionally dominated by domain-specific convolutional networks and supervised learning. We assess the performance of GPT-4o, Claude 3.5 Sonnet, and LLaMA 3.2 Vision on two benchmark datasets, UTKFace and FG-NET, without any fine-tuning or task-specific adaptation. Using eight evaluation metrics, including MAE, MSE, RMSE, MAPE, MBE, $R^2$, CCC, and $\pm$5-year accuracy, we demonstrate that general-purpose LVLMs can deliver competitive performance in zero-shot settings. Our findings highlight the emergent capabilities of LVLMs for accurate biometric age estimation and position these models as promising tools for real-world applications. Additionally, we highlight performance disparities linked to image quality and demographic subgroups, underscoring the need for fairness-aware multimodal inference. This work introduces a reproducible benchmark and positions LVLMs as promising tools for real-world applications in forensic science, healthcare monitoring, and human-computer interaction. The benchmark focuses on strict zero-shot inference without fine-tuning and highlights remaining challenges related to prompt sensitivity, interpretability, computational cost, and demographic fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。