arXiv:2501.00765cs.CVcs.LG2025-01

首个支持中文手语双向交互的统一数据集与模型,实现精准手势生成与翻译。

Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility

  • 构建首个中文手语双向数据集,含1.5万句对与8643个词汇的骨骼关键点。
  • 提出AuraLLM模型,直接评估手势姿态,实现可控协调的手语视频生成。
  • SignMST-C在PHOENIX2014-T上达32.08的BLEU-4,推动手语翻译新基准。

全球7200万听障人士主要依赖手语沟通,亟需高效的双向手语生成与翻译系统。然而,功能完整的双向系统需统一语言环境,受限于缺乏合适的统一数据集,尤其缺少用于准确手语生成(SLP)评估的姿态信息。当前的评估方法如回译忽略姿态精度,高质量协同生成仍具挑战。为此,我们推出CNText2Sign与CNSign,构成首个支持中文手语双向可访问性的统一数据集:CNText2Sign包含15,000条自然语言到手语的映射,以及8,643个词汇的标准骨骼关键点,支持姿态评估。基于此,我们提出AuraLLM模型,采用解耦架构结合姿态数据,实现新型直接手势精度评估;通过检索增强与级联词汇解析处理语义映射与未登录词,并通过姿态条件化视频合成实现全场景可控手势与面部表情协同生成。同时,手语翻译模型SignMST-C采用定向自监督预训练捕捉动态特征,在PHOENIX2014-T上达到32.08的最高BLEU-4分数。AuraLLM在直接评估下于CNText2Sign上取得50.41的BLEU-4得分,建立强性能基线。

原文摘要 · Abstract (English)

Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Language Translation systems. However, functional bidirectional systems require a unified linguistic environment, hindered by the lack of suitable unified datasets, particularly those providing the necessary pose information for accurate Sign Language Production (SLP) evaluation. Concurrently, current SLP evaluation methods like back-translation ignore pose accuracy, and high-quality coordinated generation remains challenging. To create this crucial environment and overcome these challenges, we introduce CNText2Sign and CNSign, which together constitute the first unified dataset aimed at supporting bidirectional accessibility systems for Chinese sign language; CNText2Sign provides 15,000 natural language-to-sign mappings and standardized skeletal keypoints for 8,643 vocabulary items supporting pose assessment. Building upon this foundation, we propose the AuraLLM model, which leverages a decoupled architecture with CNText2Sign's pose data for novel direct gesture accuracy assessment. The model employs retrieval augmentation and Cascading Vocabulary Resolution to handle semantic mapping and out-of-vocabulary words and achieves all-scenario production with controllable coordination of gestures and facial expressions via pose-conditioned video synthesis. Concurrently, our Sign Language Translation model SignMST-C employs targeted self-supervised pretraining for dynamic feature capture, achieving new SOTA results on PHOENIX2014-T with BLEU-4 scores up to 32.08. AuraLLM establishes a strong performance baseline on CNText2Sign with a BLEU-4 score of 50.41 under direct evaluation.

手语生成姿态评估双向翻译中文手语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。