将语音互动的预测框架移植到手语交互,验证其可行性与局限性。
Toward Signing Activity Projection in Sign Language Interaction

- 用姿态特征提取手、眼、口区域信号,构建手语活动流预测模型。
- 手部线索对转向预测有效,但整体转向预测仍困难,准确率未达理想水平。
- 为手语机器人交互提供新思路,适合具身智能与无障碍交互研究者。
社交机器人需在非语音主导的交互中保持鲁棒性,尤其面对依赖手语沟通的用户。当前一个重要能力缺口是对手语用户的预测性换轮。尽管语音活动预测(VAP)在口语交互中表现成功,但其能否迁移至手语交互尚不明确。本文首次开展将VAP架构应用于双人手语交互的迁移研究。基于Public DGS语料库的交互记录,从词汇级手语标注中提取二值化手语活动流,并构建换轮预测代理任务。模型利用每名手语者的手部、眼部及口部区域的姿态特征。结果显示,对于手语活动的持续(HOLD)和切换(SHIFT)预测具有一定前景,尤其是依赖手部线索时;然而SHIFT预测仍具挑战。这些发现初步表明,将语音交互中的预测模型迁移至手语交互具有潜力,但也存在显著局限。未来预测建模需基于手语特有事件定义,而非仅沿用语音衍生的分类体系。
原文摘要 · Abstract (English)
Social robots must interact robustly not only with users assumed by speech-centered systems but also with diverse users whose communication relies on different modalities, e.g., sign language. One important capability gap is predictive turn-taking with signing users. Although Voice Activity Projection (VAP) has been successfully used to model future voice activity in spoken interaction, it remains unclear whether the framework transfers to sign language interaction. This paper presents an initial transfer study of adapting a VAP architecture to dyadic sign language interaction. Using interaction recordings from the Public DGS Corpus, we derive binary signing activity streams from lexical sign annotations and formulate proxy tasks for turn-taking prediction. The model uses pose-derived hand, eye-region, and mouth-region features extracted for each signer. The results show that SHIFT/HOLD prediction is promising, especially with hand cues, while SHIFT-prediction remains difficult. These findings provide initial evidence for both the promise and the current limitations of transferring predictive turn-taking models from spoken interaction to sign language interaction. Predictive modeling of sign language interaction still requires sign-language-specific event definitions that go beyond speech-derived categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。