让聋人可编辑的实时手语生成系统,提升自然度与用户控制力。
Human-Centered Editable Speech-to-Sign-Language Generation via Streaming Conformer-Transformer and Resampling Hook
- 用流式Conformer+Transformer-MDN生成同步手部与面部动作
- 通过可编辑的JSON中间表示实现逐段修改,延迟仅103毫秒
- 支持聋人用户参与优化,显著提升易用性与信任感
现有端到端手语动画系统存在自然度低、表情肢体表现力弱、无法用户操控等问题。本文提出一种以人为中心的实时语音转手语动画框架,结合(1)流式Conformer编码器与自回归Transformer-MDN解码器,实现上身与面部动作的同步生成;(2)透明可编辑的JSON中间表示,使聋人用户与专家能逐段查看并修改手语内容;(3)人机协同优化循环,根据用户编辑与评分持续改进模型。系统部署于Unity3D,平均帧推理时间13毫秒,端到端延迟103毫秒(RTX 4070)。关键贡献包括面向细粒度个性化设计的以JSON为中心的编辑机制,以及首次将基于MDN的反馈环用于模型持续适配。在20名聋人使用者和5位专业译员的测试中,系统在用户体验评分(SUS)上提升13分,认知负荷降低6.7分,自然度与可信度显著优于基线(p < .001)。该工作建立了可扩展、可解释的可访问手语技术通用范式。
原文摘要 · Abstract (English)
Existing end-to-end sign-language animation systems suffer from low naturalness, limited facial/body expressivity, and no user control. We propose a human-centered, real-time speech-to-sign animation framework that integrates (1) a streaming Conformer encoder with an autoregressive Transformer-MDN decoder for synchronized upper-body and facial motion generation, (2) a transparent, editable JSON intermediate representation empowering deaf users and experts to inspect and modify each sign segment, and (3) a human-in-the-loop optimization loop that refines the model based on user edits and ratings. Deployed on Unity3D, our system achieves a 13 ms average frame-inference time and a 103 ms end-to-end latency on an RTX 4070. Our key contributions include the design of a JSON-centric editing mechanism for fine-grained sign-level personalization and the first application of an MDN-based feedback loop for continuous model adaptation. This combination establishes a generalizable, explainable AI paradigm for user-adaptive, low-latency multimodal systems. In studies with 20 deaf signers and 5 professional interpreters, we observe a +13 point SUS improvement, 6.7 point reduction in cognitive load, and significant gains in naturalness and trust (p $<$ .001) over baselines. This work establishes a scalable, explainable AI paradigm for accessible sign-language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。