arXiv:2605.14705cs.CV2026-05

用孤立手势构建连续手语对话,实现无需文字的自然手语交互。

Towards Continuous Sign Language Conversation from Isolated Signs

论文配图:Towards Continuous Sign Language Conversation from Isolated Signs
图 1 · 摘自论文原文
  • 将孤立手势重组为连贯手语对话,解决数据稀缺问题。
  • 提出BRAID模型,实现动作时长对齐与边界补全,提升流畅性。
  • 首个端到端手语对话生成模型,适合听障人士的视觉化交流需求。

手语是许多聋哑人士的主要语言,但现有对话型AI系统仍以口语或书面语为中心,限制了手语使用者的接入。为此,本文通过大规模标注的孤立手势片段构建连续手语对话:从现有对话语料中提取手语有序表达,利用这些孤立手势作为词汇级运动基元进行重组。提出SignaVox-W,目前最大的标注孤立手势词库;以及基于其构建的SignaVox-U连续3D手语对话数据集。为弥合口语与手语结构差异,采用检索引导的语音转词翻译器;为整合独立采集的手势片段,提出BRAID——一种扩散式Transformer模型,实现时长对齐与共音边界修复。基于此数据训练出SignaVox模型,可直接根据前序手语上下文生成3D身体、手部和面部动作响应,推理时不依赖文本或外部词注释。定量与定性评估表明,该方法显著提升孤立到连续动作质量,增强语义一致性,并支持可扩展的以手语者为中心的交互,更契合视觉空间表达需求。

原文摘要 · Abstract (English)

Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written language. This spoken-language-centered interface can limit access for signers for whom spoken or written language is not the most accessible medium, motivating direct sign-to-sign conversational modeling. However, sentence-level sign video data are expensive to collect and annotate, leaving existing sign translation and production models with limited vocabulary coverage and weak open-domain generalization. We address this bottleneck by constructing continuous sign conversations from isolated signs: large-scale labeled isolated clips are collected as lexically grounded motion primitives and recomposed into sign-language-ordered utterances derived from existing dialogue corpora. We introduce SignaVox-W, which provides, to our knowledge, the largest labeled isolated-sign vocabulary to date, and SignaVox-U, a continuous 3D sign conversation dataset built from SignaVox-W. To bridge structural mismatch between spoken and signed languages, we use a retrieval-guided spoken-to-gloss translator; to bridge independently collected isolated clips, we propose BRAID, a diffusion Transformer that performs duration alignment and co-articulatory boundary inpainting. With the resulting data, we train SignaVox, a direct sign-to-sign conversational model that generates 3D body, hand, and facial motion responses from prior signing context without spoken-language text or externally provided glosses at inference time. Quantitative and qualitative evaluations show improved isolated-to-continuous motion quality, stronger response-level semantic alignment, and scalable signer-centered interaction that better supports visual-spatial articulation.

手语生成3D动作对话系统聋哑交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。