arXiv:2604.18184cs.CV2026-04

通过正视图引导多视角手语识别,提升真实场景下鲁棒性。

CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition

论文配图:CanonSLR: Canonical-View Guided Multi-View Continuous Sign Language Recognition
图 1 · 摘自论文原文
  • 用正视图数据训练教师网络,为多视角学生网络提供时序监督。
  • 在七个视角上性能超越现有方法,非正面视角识别准确率显著提升。
  • 适合做多视角手语识别、无障碍交互系统研发的团队参考。

连续手语识别近年进展显著,但多数方法基于单视角设定,在真实场景中对视角变化仍不够鲁棒。为此,我们提出CanonSLR框架,采用正视图锚定的师生学习策略:由正视图训练的教师网络为全视角学生网络提供标准时序监督。为减少跨视角语义差异,提出序列级软目标蒸馏,将正视图结构化时序知识迁移到非正视样本,缓解遮挡与投影变化导致的词边界模糊和类别混淆。同时引入时间运动关系增强模块,显式建模高层视觉特征中的运动感知时序关系,强化稳定动态表征并抑制视角敏感的外观干扰。为支持多视角研究,构建通用多视角手语视频生成管道,将原始单视角RGB视频转换为语义一致、时序连贯、视角可控的多视角视频。基于此,扩展PHOENIX-2014T和CSL-Daily为七视角基准PT14-MV与CSL-MV,为多视角手语识别提供新实验基础。大量实验证明,CanonSLR在多视角设置下持续优于现有方法,尤其在挑战性的非正面视角表现更优。

原文摘要 · Abstract (English)

Continuous Sign Language Recognition (CSLR) has achieved remarkable progress in recent years; however, most existing methods are developed under single-view settings and thus remain insufficiently robust to viewpoint variations in real-world scenarios. To address this limitation, we propose CanonSLR, a canonical-view guided framework for multi-view CSLR. Specifically, we introduce a frontal-view-anchored teacher-student learning strategy, in which a teacher network trained on frontal-view data provides canonical temporal supervision for a student network trained on all viewpoints. To further reduce cross-view semantic discrepancy, we propose Sequence-Level Soft-Target Distillation, which transfers structured temporal knowledge from the frontal view to non-frontal samples, thereby alleviating gloss boundary ambiguity and category confusion caused by occlusion and projection variation. In addition, we introduce Temporal Motion Relational Enhancement to explicitly model motion-aware temporal relations in high-level visual features, strengthening stable dynamic representations while suppressing viewpoint-sensitive appearance disturbances. To support multi-view CSLR research, we further develop a universal multi-view sign language data construction pipeline that transforms original single-view RGB videos into semantically consistent, temporally coherent, and viewpoint-controllable multi-view sign language videos. Based on this pipeline, we extend PHOENIX-2014T and CSL-Daily into two seven-view benchmarks, namely PT14-MV and CSL-MV, providing a new experimental foundation for multi-view CSLR. Extensive experiments on PT14-MV and CSL-MV demonstrate that CanonSLR consistently outperforms existing approaches under multi-view settings and exhibits stronger robustness, especially on challenging non-frontal views.

手语识别多视角蒸馏视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。