arXiv:2605.01720cs.CVcs.AI2026-05

构建了200万片段的多语言手语姿态数据集,支持真实场景下的手语识别与生成。

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages

论文配图:SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
图 1 · 摘自论文原文
  • 将多源手语视频统一转为2D姿态序列,适配现代姿态驱动模型
  • 覆盖55种手语,含两百万个视频片段,保留真实拍摄条件
  • 适合手语识别、生成及跨语言建模研究者使用

现有大规模手语资源通常仅提供原始视频与文本的对齐标注,且多在实验室环境下采集。这类资源虽有助于语义理解,却难以直接对接开放世界识别与翻译,或现代姿态驱动的手语视频生成框架:1. 基于RGB的预训练识别模型对背景和着装依赖强,在开放场景中鲁棒性较差;2. 当前姿态引导的图像/视频生成模型多采用如DWPose等统一关键点表示作为控制接口。目前手语领域仍缺乏可直接对接此姿态原生范式的资源。我们提出SignVerse-2M,一个大规模多语言姿态原生手语数据集,用于手语姿态建模与评估。该数据集基于公开多语言手语视频资源,采用统一预处理流程应用DWPose,将原始视频转换为可直接用于建模的2D姿态序列,形成约两百万个片段的综合语料库,涵盖55种以上手语。相比实验室数据集,该资源保留真实视频的拍摄条件与说话人多样性,同时通过统一姿态表示减少外观差异。为此,我们进一步提供数据构建流程、任务定义及一个简单的SignDW Transformer基线模型,验证该资源在多语言姿态空间建模中的可行性及其与现代姿态驱动流程的兼容性,并讨论其可支持的评估命题与当前局限。

原文摘要 · Abstract (English)

Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do not directly provide a unified interface for open-world recognition and translation, or for modern pose-driven sign language video generation frameworks: 1. RGB-based pretrained recognition models depend heavily on fixed backgrounds or clothing conditions during recording, and are less robust in open-world settings than style-agnostic pose-processing models. 2. Recent pose-guided image/video generation models mostly use a unified keypoint representation such as DWPose as their control interface. At present, the sign language field still lacks a data resource that can directly interface with this modern pose-native paradigm while also targeting real-world open scenarios. We present SignVerse-2M, a large-scale multilingual pose-native dataset for sign language pose modeling and evaluation. Built from publicly available multilingual sign language video resources, it applies DWPose in a unified preprocessing pipeline to convert raw videos into 2D pose sequences that can be used directly for modeling, resulting in a consolidated corpus of about two million clips covering more than 55 sign languages. Unlike many laboratory datasets, this resource preserves the recording conditions and speaker diversity of real-world videos while reducing appearance variation through a unified pose representation. Toward this goal, we further provide the data construction pipeline, task definitions, and a simple SignDW Transformer baseline, demonstrating the feasibility of this resource for multilingual pose-space modeling and its compatibility with modern pose-driven pipelines, while discussing the evaluation claims it can support as well as its current limitations.

手语识别姿态生成多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。