用激活调控让视觉模型不再靠认脸猜年龄,提升真实年龄预测准确率。
When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering
- 通过干预视觉语言模型隐藏状态,打断‘认脸猜年龄’的捷径策略。
- 在多个基准上将平均绝对误差降低最多25%,对普通人和未见身份均有效。
- 适合关注模型公平性与真实场景鲁棒性的研究人员参考。
为保护未成年人免受有害内容影响,自动化年龄估计至关重要,当前基于视觉-语言模型(VLMs)的方法表现领先。然而我们发现,零样本特性导致一种名为‘身份捷径’的现象:模型并非从视觉特征推断年龄,而是通过识别人物身份,从记忆中获取其年龄。这使得非名人被误判为名人时产生严重错误预测。同时,在以名人图像为主导的基准上,模型表现出虚假的高噪声与对抗扰动鲁棒性。为此,我们提出一种激活调控方法,通过干预VLM的隐藏状态来抑制该捷径行为。该方法显著提升对已知与未知身份的年龄估计精度,在主流基准上平均绝对误差最高降低25%。
原文摘要 · Abstract (English)
Different age-related regulations have been proposed to protect minors from harmful content and interactions online. Automated age estimation is central to enforcing such regulations, and vision-language models (VLMs) achieve state-of-the-art performance on this task. However, we find that the zero-shot nature of VLM-based age estimation produces an unexpected side effect we call the identity shortcut: Instead of estimating age from visual features, VLMs tend to identify the depicted person and infer their age from memorized knowledge. This phenomenon leads to substantially incorrect predictions when non-celebrities are misidentified as celebrities. It also produces deceptively high robustness to noise and adversarial perturbations on celebrity images, which dominate popular benchmarks. To mitigate this, we propose an activation steering method that suppresses the shortcut by intervening on the hidden states of the VLM. This method improves age estimation accuracy for both memorized and unseen identities, reducing mean absolute error by up to 25% across popular benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。