首个中文手语理解基准,评估多模态大模型在手语识别中的表现。
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

- 构建基于国家标准的手语数据集,确保语义一致性。
- 覆盖视频、图像、文本三模态,支持细粒度手语分析。
- 发现当前模型仍远低于人类水平,尤其在不同手语形式上表现不一。
由于大语言模型(LLMs)的发展,手语研究取得了显著进展。然而,大模型在多模态环境下理解手语的内在能力仍缺乏深入探索。为此,我们提出CNSL-bench,首个面向中文国家手语的多模态大语言模型(MLLMs)评估基准。该基准具有三大特点:1)权威性基础,基于官方标准《国家通用手语词典》,避免地区或非规范变体带来的歧义,确保语义定义一致;2)多模态覆盖,提供对齐的文本描述、示例图像与手语视频;3)动作多样性,支持对关键手动表达形式(如空中写字、手指拼写、汉语手语字母)的细粒度分析。利用CNSL-bench,我们全面评估了21个开源及专有最新版MLLMs。结果表明,尽管多模态建模持续进步,当前模型仍显著落后于人类表现,在不同输入模态和手语形式间存在系统性差异。诊断分析进一步显示,部分性能瓶颈并未随推理能力提升而缓解,且指令遵循鲁棒性在不同模型间差异显著。
原文摘要 · Abstract (English)
Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, especially in multimodal contexts, remains underexplored. To address this limitation, we introduce CNSL-bench, the first comprehensive Chinese em{National Sign Language benchmark designed for evaluating multimodal large language models (MLLMs) in sign language understanding. The proposed CNSL-bench is characterized by: 1) Authoritative grounding, as it is anchored to the officially standardized \textit{National Common Sign Language Dictionary, mitigating ambiguity from regional or non-canonical variants and ensuring consistent semantic definitions; 2) Multimodal coverage, providing aligned textual descriptions, illustrative images, and sign language videos; and 3) Articulatory diversity, supporting fine-grained analysis across key manual articulatory forms, including air-writing, finger-spelling, and the Chinese manual-alphabet. Using CNSL-bench, we extensively evaluate 21 open-source and proprietary up-to-date MLLMs. Our results reveal that, despite recent advances in multimodal modeling, current MLLMs remain substantially inferior to human performance, exhibiting systematic disparities across input modalities and manual articulatory forms. Additional diagnostic analyses suggest that several performance limitations persist beyond improvements in reasoning and that instruction-following robustness varies substantially across models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。