arXiv:2608.25493cs.CV2026-08中稿 · 37th British Machi…

用多模态大模型引导时序对齐,统一手语识别与定位任务。

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

论文配图:SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting
图 1 · 摘自论文原文
  • 利用多模态大模型生成运动描述作为辅助语义线索,实现小批量稳定对齐。
  • 引入多尺度时序适配器提升时序特征学习,结合边界感知定位模块增强定位精度。
  • 在四个数据集上同时提升识别与定位性能,适合手语理解与人机交互研究者。

连续手语识别(CSLR)旨在弱序列级监督下从非分割手语视频中识别词序列。现有方法依赖句子级词标注,难以提供细粒度的时空与语义指导。传统视频-文本对齐需大规模批量训练,内存开销大。本文提出SMART,一种基于多模态大模型(MLLM)引导的时序对齐框架,可联合完成手语识别与定位。SMART利用MLLM生成的运动描述作为辅助语义线索,在小批量训练下实现稳定视频-文本对齐。为提升时序表示学习,设计了多尺度时序适配器以建模变换器编码过程中的时序交互。针对密集时序定位,引入CSFormer——一个由CSLR引导的定位模块,将识别获得的词证据注入边界感知定位网络。该统一框架使识别特征惠及定位,定位监督亦反哺弱监督的CTC识别。在PHOENIX14-T、CSL-Daily、Large-scale KSL和Disaster and Safety KSL共四个手语基准上验证,SMART在识别与定位任务中均表现优异。

原文摘要 · Abstract (English)

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional video-text alignment also requires large batch sizes, making it inefficient for memory-intensive sign language video training. In this work, we propose SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting. SMART uses MLLMgenerated motion descriptions as auxiliary semantic cues and performs stable videotext alignment under small-batch training. To improve temporal representation learning, we introduce a Multi-Scale Temporal Adapter that models temporal interactions during transformer encoding. For dense temporal localization, SMART incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network. This unified framework enables CSLR features to benefit spotting, while spotting supervision complements weak CTC-based recognition. Experiments on four sign language benchmarks, including PHOENIX14-T, CSL-Daily, Large-scale KSL, and Disaster and Safety KSL datasets, demonstrate the effectiveness of SMART across both recognition and spotting tasks.

手语识别时序对齐多模态定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。