用注意力机制分析手语中面部表情的作用
Cross-Attention Based Influence Model for Manual and Nonmanual Sign Language Analysis
- 双流编码器分别处理面部和上半身动作
- 提出并行交叉注意力解码,量化各模态贡献度
- 首次定量揭示面部表情对手语理解的关键影响
美式手语(ASL)的完整语义不仅依赖手势,还依赖面部表情等非手动标记(NMM)。尽管已有研究致力于手语到口语/书面语的理解,但多数仅关注手势特征。本文采用先进的神经机器翻译方法,系统考察面部表情对手语短语理解的贡献程度。提出一种双流编码器架构,分别处理面部和上半身(含手部)动作;设计新的并行交叉注意力解码机制,使两个编码流同时输入解码器的不同注意力堆栈,通过分析其注意力权重,定量评估面部、身体与手势在翻译任务中的相对重要性。
原文摘要 · Abstract (English)
Both manual (relating to the use of hands) and non-manual markers (NMM), such as facial expressions or mouthing cues, are important for providing the complete meaning of phrases in American Sign Language (ASL). Efforts have been made in advancing sign language to spoken/written language understanding, but most of these have primarily focused on manual features only. In this work, using advanced neural machine translation methods, we examine and report on the extent to which facial expressions contribute to understanding sign language phrases. We present a sign language translation architecture consisting of two-stream encoders, with one encoder handling the face and the other handling the upper body (with hands). We propose a new parallel cross-attention decoding mechanism that is useful for quantifying the influence of each input modality on the output. The two streams from the encoder are directed simultaneously to different attention stacks in the decoder. Examining the properties of the parallel cross-attention weights allows us to analyze the importance of facial markers compared to body and hand features during a translating task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。