arXiv:2507.03703cs.CVcs.AI2025-07中稿 · the international …被引 1

用大模型解决手语识别中的歧义问题,不需训练就能提升准确率。

Sign Spotting Disambiguation using Large Language Models

  • 通过大模型进行上下文感知的词义消歧,无需微调。
  • 在真实和合成数据集上准确率显著优于传统方法。
  • 适用于手语标注、翻译系统,尤其适合资源稀缺场景。

手语识别任务旨在从连续手语视频中定位并识别单个手势,对大规模数据标注和缓解手语翻译中的数据稀缺问题至关重要。尽管自动手语识别具有实现帧级监督的巨大潜力,但仍面临词汇僵化和连续手势流中固有歧义等挑战。为此,我们提出一种全新的、无需训练的框架,利用大语言模型(LLM)显著提升手语识别质量。该方法提取全局时空与手形特征,通过动态时间规整和余弦相似度与大规模手语词典进行匹配,这种基于词典的匹配天然具备词汇灵活性,且无需模型重训练。为降低匹配过程中的噪声与歧义,大语言模型通过束搜索执行上下文感知的词义消歧,同样无需微调。在合成与真实世界手语数据集上的大量实验表明,该方法在准确率和句子流畅性方面均优于传统方法,凸显了大语言模型在推进手语识别方面的潜力。

原文摘要 · Abstract (English)

Sign spotting, the task of identifying and localizing individual signs within continuous sign language video, plays a pivotal role in scaling dataset annotations and addressing the severe data scarcity issue in sign language translation. While automatic sign spotting holds great promise for enabling frame-level supervision at scale, it grapples with challenges such as vocabulary inflexibility and ambiguity inherent in continuous sign streams. Hence, we introduce a novel, training-free framework that integrates Large Language Models (LLMs) to significantly enhance sign spotting quality. Our approach extracts global spatio-temporal and hand shape features, which are then matched against a large-scale sign dictionary using dynamic time warping and cosine similarity. This dictionary-based matching inherently offers superior vocabulary flexibility without requiring model retraining. To mitigate noise and ambiguity from the matching process, an LLM performs context-aware gloss disambiguation via beam search, notably without fine-tuning. Extensive experiments on both synthetic and real-world sign language datasets demonstrate our method's superior accuracy and sentence fluency compared to traditional approaches, highlighting the potential of LLMs in advancing sign spotting.

手语识别大模型应用词义消歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。