构建了12种手语的视频语料库,支持跨语言研究与技术开发。
The TUB Sign Language Corpus Collection
- 从新闻、政府和教育频道采集多源视频,构建平行语料库
- 总时长超1300小时,含1400万词符的多语言字幕
- 首次提供8种拉美手语语料,德语手语数据量扩大十倍
我们发布了一个包含12种手语的并行语料库,以视频形式呈现,并配有对应国家主流口语的字幕。整个语料库包含超过1300小时的视频内容,共4,381个视频文件,配套130万条字幕,总计1400万词符。特别值得注意的是,该语料库首次为8种拉丁美洲手语提供了系统性的并行数据,而德国手语语料规模是此前可用数据的十倍。数据通过收集和处理来自多个在线来源的视频(主要是新闻节目、政府机构及教育频道)获得,包括数据采集、内容创作者沟通、使用授权获取、网页抓取和视频裁剪等流程。本文还提供了语料库的统计数据及数据收集方法概述。
原文摘要 · Abstract (English)
We present a collection of parallel corpora of 12 sign languages in video format, together with subtitles in the dominant spoken languages of the corresponding countries. The entire collection includes more than 1,300 hours in 4,381 video files, accompanied by 1,3~M subtitles containing 14~M tokens. Most notably, it includes the first consistent parallel corpora for 8 Latin American sign languages, whereas the size of the German Sign Language corpora is ten times the size of the previously available corpora. The collection was created by collecting and processing videos of multiple sign languages from various online sources, mainly broadcast material of news shows, governmental bodies and educational channels. The preparation involved several stages, including data collection, informing the content creators and seeking usage approvals, scraping, and cropping. The paper provides statistics on the collection and an overview of the methods used to collect the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。