arXiv:2409.11960cs.CV2024-09被引 6

构建首个复杂场景中文连续手语数据集,提升真实环境识别能力

A Chinese Continuous Sign Language Dataset Based on Complex Environments

  • 提出时频网络,融合时间与频谱信息提升特征表达
  • 在5988段真实场景视频上实现显著性能提升
  • 适合手语识别、人机交互及多模态研究者使用

当前连续手语识别(CSLR)研究的瓶颈在于,大多数公开数据集仅限于实验室环境或电视节目录制,背景单一且光照均匀,与真实场景差异显著。为解决此问题,我们构建了一个基于复杂环境的大型中文连续手语数据集,命名为复杂环境-中文手语数据集(CE-CSL)。该数据集包含5,988段日常场景中的连续手语视频,涵盖70余种不同复杂背景,以确保代表性和泛化能力。针对复杂背景对手语识别的影响,我们提出一种时频网络(TFNet)模型,通过提取帧级特征,并分别利用时间和频谱信息生成序列特征后融合,实现高效准确的连续手语识别。实验表明,该方法在CE-CSL上取得显著性能提升,验证了其在复杂背景下的有效性。此外,该方法在三个公开可用的CSL数据集上也表现出竞争力。

原文摘要 · Abstract (English)

The current bottleneck in continuous sign language recognition (CSLR) research lies in the fact that most publicly available datasets are limited to laboratory environments or television program recordings, resulting in a single background environment with uniform lighting, which significantly deviates from the diversity and complexity found in real-life scenarios. To address this challenge, we have constructed a new, large-scale dataset for Chinese continuous sign language (CSL) based on complex environments, termed the complex environment - chinese sign language dataset (CE-CSL). This dataset encompasses 5,988 continuous CSL video clips collected from daily life scenes, featuring more than 70 different complex backgrounds to ensure representativeness and generalization capability. To tackle the impact of complex backgrounds on CSLR performance, we propose a time-frequency network (TFNet) model for continuous sign language recognition. This model extracts frame-level features and then utilizes both temporal and spectral information to separately derive sequence features before fusion, aiming to achieve efficient and accurate CSLR. Experimental results demonstrate that our approach achieves significant performance improvements on the CE-CSL, validating its effectiveness under complex background conditions. Additionally, our proposed method has also yielded highly competitive results when applied to three publicly available CSL datasets.

手语识别多模态真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。