arXiv:2609.06296cs.CVcs.CL2026-09

用时间维度自蒸馏学习手语表示,提升跨任务表现

SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

论文配图:SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation
图 1 · 摘自论文原文
  • 将自监督蒸馏从图像空间转到手语时间序列,分左右手和面部流处理
  • 在多个手语识别任务上达到顶尖性能,优于现有自监督方法
  • 适合手语识别、翻译与无障碍应用研究者参考

自监督手语表征学习需建模两个非自然图像的核心特性:手势由有限解剖部位产生,且语义依赖于这些部位的时间组织。我们提出SignDino,一种将DINOv3学生-教师框架从图像补丁空间迁移至追踪手语流时间域的自监督视频编码器。每段视频通过检测-跟踪流水线(YOLOv8n+ByteTrack)分解为左手指、右手指和面部三路流。固定冻结的DINOv3 ViT-B/16对每个帧的解剖区域进行嵌入,而轻量级时间注意力网络作为学生与指数移动平均教师模型。通过时间域DINO自蒸馏、iBOT风格的帧级掩码预测、KoLeo特征扩散及帧间相似性结构的Gram锚定进行训练。该设计保持强图像级视觉原语不变,仅学习解剖部位状态随时间演化规律。在手语到英文翻译、孤立手语识别与指拼检测基准上评估,SignDino展现出强大公共自监督表征,在匹配下游评估条件下表现竞争力或达到当前最优。

原文摘要 · Abstract (English)

Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.

手语识别自监督学习时间建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。