arXiv:2507.21104cs.CLcs.AI2025-07

构建首个乌拉圭手语翻译公开数据集,助力无障碍交流

iLSU-T: an Open Dataset for Uruguayan Sign Language Translation

  • 采集185小时乌拉圭手语视频,含音视频与文字转录
  • 3种主流算法在该数据集上建立基线,验证可用性
  • 适合手语识别、多模态翻译研究者使用

近年来,自动手语翻译在计算机视觉与计算语言学领域受到广泛关注。由于各国手语具有独特性,机器翻译需依赖本地化数据以开发新技术或适配现有方法。本文提出iLSU-T,一个开放的乌拉圭手语翻译数据集,包含来自公共电视广播的超过185小时的彩色视频,配有音频与文字转录。数据涵盖多样主题,由18位专业手语翻译员参与。通过三种前沿翻译算法进行实验,旨在建立该数据集的基准并评估数据处理流程的有效性。实验凸显了本地化手语数据对提升可访问性与包容性的关键作用。数据与代码已公开。

原文摘要 · Abstract (English)

Automatic sign language translation has gained particular interest in the computer vision and computational linguistics communities in recent years. Given each sign language country particularities, machine translation requires local data to develop new techniques and adapt existing ones. This work presents iLSU T, an open dataset of interpreted Uruguayan Sign Language RGB videos with audio and text transcriptions. This type of multimodal and curated data is paramount for developing novel approaches to understand or generate tools for sign language processing. iLSU T comprises more than 185 hours of interpreted sign language videos from public TV broadcasting. It covers diverse topics and includes the participation of 18 professional interpreters of sign language. A series of experiments using three state of the art translation algorithms is presented. The aim is to establish a baseline for this dataset and evaluate its usefulness and the proposed pipeline for data processing. The experiments highlight the need for more localized datasets for sign language translation and understanding, which are critical for developing novel tools to improve accessibility and inclusion of all individuals. Our data and code can be accessed.

手语翻译多模态数据无障碍公开数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。