评测视觉语言模型对手语形义映射的直观性,发现模型仍远低于人类水平。
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
- 构建视频基准测试,评估模型对手势形态、意义透明度和图标化程度的判断。
- 模型在手势形态预测上部分达标,但意义推断能力远不及人类,图标化评分相关性弱。
- 顶尖模型与人类对形义关联的敏感度更接近,提示需加强具身学习与人类信号建模。
象似性(Iconicity)指语言形式与意义之间的相似性,在手语中普遍存在,为视觉接地提供了天然测试场景。对于视觉语言模型(VLMs),挑战在于从动态人体动作中恢复这种关键映射,而非静态上下文。我们提出视觉象似性挑战(Visual Iconicity Challenge),一个基于视频的新基准,将心理语言学测量方法适配于三个任务:(i) 语音学手势形态预测(如手形、位置),(ii) 透明度(从视觉形式推断意义),(iii) 分级象似性评分。我们在荷兰手语数据集上,以零样本和少样本设置评估13个先进VLMs,并与人类基线比较。在语音学形态预测上,模型能捕捉部分手形和位置信息,但性能仍低于人类;在透明度任务中,模型表现显著落后于人类;仅有顶级模型与人类象似性评分呈现中等相关性。有趣的是,语音学形态预测能力强的模型,其象似性判断也更接近人类,表明对视觉结构的共同敏感性。研究结果验证了这些诊断任务的有效性,推动采用以人为中心的信号与具身学习方法来建模象似性,提升多模态模型的视觉接地能力。
原文摘要 · Abstract (English)
Iconicity, the resemblance between linguistic form and meaning, is pervasive in signed languages, offering a natural testbed for visual grounding. For vision-language models (VLMs), the challenge is to recover such essential mappings from dynamic human motion rather than static context. We introduce the Visual Iconicity Challenge, a novel video-based benchmark that adapts psycholinguistic measures to evaluate VLMs on three tasks: (i) phonological sign-form prediction (e.g., handshape, location), (ii) transparency (inferring meaning from visual form), and (iii) graded iconicity ratings. We assess 13 state-of-the-art VLMs in zero- and few-shot settings on Sign Language of the Netherlands and compare them to human baselines. On phonological form prediction, VLMs recover some handshape and location detail but remain below human performance; on transparency, they are far from human baselines; and only top models correlate moderately with human iconicity ratings. Interestingly, models with stronger phonological form prediction correlate better with human iconicity judgment, indicating shared sensitivity to visually grounded structure. Our findings validate these diagnostic tasks and motivate human-centric signals and embodied learning methods for modelling iconicity and improving visual grounding in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。