用视觉语言模型自动采集社交平台手语数据,降低人工成本。
Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media
- 用VLM检测面部可见性、手势活动和视频文字,实现自动标注与过滤
- 在8种手语的TikTok数据上构建TikTok-SL-8数据集,支持德语与美式手语测试
- 为手语翻译模型提供弱监督预训练数据,适合资源有限的研究者
现有手语翻译(SLT)数据集规模小、多语言覆盖不足,且依赖专家标注与受控录制,成本高昂。近期视觉语言模型(VLMs)展现出强大的评估与实时辅助能力,但其在手语数据采集中的潜力尚未被挖掘。本文提出首个基于VLM的自动化标注与筛选框架,减少对人工的依赖,同时保证数据质量。该方法应用于八种手语的TikTok视频,并对已有的德语手语数据集YouTube-SL-25进行额外评估。框架包含面部可见性检测、手势活动识别、视频内容文本提取及视频与文本对齐验证四个步骤,实现通用的数据过滤、标注与验证。基于生成的TikTok-SL-8数据集,我们评估了两个现成的SLT模型在德语与美式手语上的表现,旨在建立基线并检验模型在自动提取、略有噪声数据上的鲁棒性。本工作实现了手语翻译的可扩展弱监督预训练,推动从社交媒体获取数据。
原文摘要 · Abstract (English)
Most existing sign language translation (SLT) datasets are limited in scale, lack multilingual coverage, and are costly to curate due to their reliance on expert annotation and controlled recording setup. Recently, Vision Language Models (VLMs) have demonstrated strong capabilities as evaluators and real-time assistants. Despite these advancements, their potential remains untapped in the context of sign language dataset acquisition. To bridge this gap, we introduce the first automated annotation and filtering framework that utilizes VLMs to reduce reliance on manual effort while preserving data quality. Our method is applied to TikTok videos across eight sign languages and to the already curated YouTube-SL-25 dataset in German Sign Language for the purpose of additional evaluation. Our VLM-based pipeline includes a face visibility detection, a sign activity recognition, a text extraction from video content, and a judgment step to validate alignment between video and text, implementing generic filtering, annotation and validation steps. Using the resulting corpus, TikTok-SL-8, we assess the performance of two off-the-shelf SLT models on our filtered dataset for German and American Sign Languages, with the goal of establishing baselines and evaluating the robustness of recent models on automatically extracted, slightly noisy data. Our work enables scalable, weakly supervised pretraining for SLT and facilitates data acquisition from social media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。