arXiv:2510.13364cs.CVcs.AI2025-10被引 1

用语言当标签,让模型零样本识别日常姿势,发现越简单提示效果越好。

Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity

  • 用三种提示层级测试视觉相似姿势分类,从简单到详细
  • 顶尖模型用最简提示准确率达68.8%,加描述后跌至55.1%
  • 低性能模型反而因具体身体线索提示表现更好,适合小数据场景

近期视觉-语言模型(VLMs)通过在共享空间中对齐图像与文本,实现零样本分类,适用于数据稀缺场景。然而,提示设计对识别视觉相似类别(如人体姿态)的影响尚不明确。本研究在仅285张图的COCO衍生数据集上,考察了提示具体性对坐、站、走/跑三类姿态零样本分类的影响。评估了OpenCLIP、MetaCLIP 2和SigLip等现代VLMs,采用三级递进式提示设计。结果揭示出反直觉趋势:对表现最佳的模型(MetaCLIP 2和OpenCLIP),最基础提示始终最优;添加描述性细节显著降低性能——例如MetaCLIP 2多类准确率从68.8%降至55.1%,此现象称为“提示过拟合”。相反,性能较低的SigLip模型在加入更多基于身体线索的详细提示后,对模糊类别分类有所提升。

原文摘要 · Abstract (English)

Recent Vision-Language Models (VLMs) enable zero-shot classification by aligning images and text in a shared space, a promising approach for data-scarce conditions. However, the influence of prompt design on recognizing visually similar categories, such as human postures, is not well understood. This study investigates how prompt specificity affects the zero-shot classification of sitting, standing, and walking/running on a small, 285-image COCO-derived dataset. A suite of modern VLMs, including OpenCLIP, MetaCLIP 2, and SigLip, were evaluated using a three-tiered prompt design that systematically increases linguistic detail. Our findings reveal a compelling, counter-intuitive trend: for the highest-performing models (MetaCLIP 2 and OpenCLIP), the simplest, most basic prompts consistently achieve the best results. Adding descriptive detail significantly degrades performance for instance, MetaCLIP 2's multi-class accuracy drops from 68.8\% to 55.1\% a phenomenon we term "prompt overfitting". Conversely, the lower-performing SigLip model shows improved classification on ambiguous classes when given more descriptive, body-cue-based prompts.

零样本视觉语言模型姿态识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。