构建百万级多语种手语数据集,提升模型在真实场景下的泛化能力
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

- 通过3D高斯溅射和扩散模型生成视角、背景与签者多样性
- 在跨视角、跨身份等条件下显著提升模型鲁棒性,保持原有性能
- 开源数据集与基准测试,适合手语识别与翻译研究者使用
手语模型通常在受控环境下训练,缺乏视角、背景和签者多样性,导致实际应用中泛化能力差。本文提出SignNet-1M,一个涵盖美式手语(ASL)、中国手语(CSL)和德语手语(DGS)的大规模增强数据集。该数据集通过三种方式模拟真实变化:(i) 利用3D高斯溅射(3DGS)实现新视角渲染(旋转与缩放);(ii) 使用扩散模型替换背景与签者,同时保留手势动作与语言内容;(iii) 添加姿态/时间扰动及视频压缩伪影等后渲染增强,更贴近野外录制效果。除数据发布外,还提供统一下游任务基准(如翻译与识别),并进行各增强组件的消融实验。实验表明,使用SignNet-1M训练可显著提升模型在跨视角、跨背景、跨身份及后处理干扰下的泛化能力,同时保持良好分布内性能。数据集、完整增强流程与基准已开放:https://signnet.chatsign.ai/
原文摘要 · Abstract (English)
Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity diversity, leading to poor robustness under real-world distribution shifts. We introduce SignNet-1M, a large-scale augmented dataset spanning ASL, CSL, and German Sign Language (DGS). SignNet-1M synthesizes realistic variations along three axes: (i) novel-view rendering (rotation and zoom) via 3D Gaussian Splatting (3DGS), (ii) scene/identity editing via diffusion models for background replacement and signer substitution while preserving sign motion and linguistic content, and (iii) post-rendering augmentations that emulate capture and compression artifacts (e.g., pose/temporal perturbations and video-level corruptions) to better match in-the-wild recordings. Beyond data release, we provide a unified benchmark suite across downstream tasks (e.g., translation and recognition) and ablations that isolate each augmentation component. Experiments across backbones show that training with SignNet-1M consistently improves generalization under cross-view, cross-background, cross-identity, and post-rendering shifts, while maintaining strong in-distribution performance. The dataset, full augmentation pipeline, and benchmark are available at https://signnet.chatsign.ai/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。