arXiv:2503.03287cs.CV2025-03被引 7

用语法规则提升手语视频字幕对齐精度

Deep Understanding of Sign Language for Sign to Subtitle Alignment

  • 利用英国手语语法预处理字幕,增强语义一致性
  • 仅在手语实际出现时预测时间位置,提升定位准确率
  • 自训练生成更准伪标签,适合小样本场景

本研究旨在解决手语视频中异步字幕对齐问题,尤其在标注数据有限的情况下。提出新框架:(1) 利用英国手语基本语法规则预处理输入字幕;(2) 设计选择性对齐损失,仅在手语实际出现时优化时间定位;(3) 通过精炼伪标签的自训练策略,提升标签准确性。实验表明,该方法在帧级准确率和F1分数上均显著超越现有基线,有效提升手语视频与字幕的对齐效果,具备手语翻译应用潜力,尤其适用于大规模人工标注困难的场景。

原文摘要 · Abstract (English)

The objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data. To achieve this goal, we propose a novel framework with the following contributions: (1) we leverage fundamental grammatical rules of British Sign Language (BSL) to pre-process the input subtitles, (2) we design a selective alignment loss to optimise the model for predicting the temporal location of signs only when the queried sign actually occurs in a scene, and (3) we conduct self-training with refined pseudo-labels which are more accurate than the heuristic audio-aligned labels. From this, our model not only better understands the correlation between the text and the signs, but also holds potential for application in the translation of sign languages, particularly in scenarios where manual labelling of large-scale sign data is impractical or challenging. Extensive experimental results demonstrate that our approach achieves state-of-the-art results, surpassing previous baselines by substantial margins in terms of both frame-level accuracy and F1-score. This highlights the effectiveness and practicality of our framework in advancing the field of sign language video alignment and translation.

手语识别字幕对齐小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。