arXiv:2410.14506cs.CLcs.AI2024-10中稿 · NeurIPS被引 1

分析视觉手语翻译模型如何关注视频帧与词符的对齐关系。

SignAttention: On the Interpretability of Transformer Models for Sign Language Translation

  • 通过注意力机制解析手语视频到词符的映射过程。
  • 发现模型关注帧组而非单帧,且对齐模式随词符增多减弱。
  • 揭示模型从依赖视频到依赖已生成词符的注意力转移规律。

本文首次对基于Transformer的手语翻译(SLT)模型进行系统性可解释性分析,聚焦于从希腊手语视频到词符和文本的翻译任务。基于希腊手语数据集,我们研究了模型中的注意力机制,以理解其如何将视觉输入与序列化词符对齐。分析表明,模型关注的是帧的聚类而非单个帧,词符与姿势之间呈现出对角线对齐模式,但随着词符数量增加,该模式逐渐模糊。我们还考察了每个解码步骤中交叉注意力与自注意力的相对贡献,发现模型初期依赖视频帧,但随着翻译推进,逐渐转向关注先前生成的词符。该研究深化了对手语翻译模型运作机制的理解,为构建更透明、可靠的实时翻译系统奠定基础。

原文摘要 · Abstract (English)

This paper presents the first comprehensive interpretability analysis of a Transformer-based Sign Language Translation (SLT) model, focusing on the translation from video-based Greek Sign Language to glosses and text. Leveraging the Greek Sign Language Dataset, we examine the attention mechanisms within the model to understand how it processes and aligns visual input with sequential glosses. Our analysis reveals that the model pays attention to clusters of frames rather than individual ones, with a diagonal alignment pattern emerging between poses and glosses, which becomes less distinct as the number of glosses increases. We also explore the relative contributions of cross-attention and self-attention at each decoding step, finding that the model initially relies on video frames but shifts its focus to previously predicted tokens as the translation progresses. This work contributes to a deeper understanding of SLT models, paving the way for the development of more transparent and reliable translation systems essential for real-world applications.

手语翻译注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。