arXiv:2412.16944cs.CVcs.MM2024-12中稿 · ICASSP 2025被引 16

提升手语生成的语义一致性,让手语动作与语言更匹配。

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

  • 用Transformer建模语言与手势的细粒度顺序对齐
  • 在PHOENIX14T上超越现有最佳方法,提升生成质量
  • 适合手语合成、跨模态对齐研究者参考

手语生成(SLP)旨在将口语句子转换为对应的手语视频,其中符号到动作的转换(G2P)是关键步骤。由于语言与视觉间的语义鸿沟以及缺乏强监督的词-动作对应标签,手语生成面临语言-视觉一致性难题。本文提出基于Transformer的語言-視覺單調一致網絡(LVMCN),通过交叉模态語義對齊器(CSA)和多模態語義比較器(MSC),分別約束細粒度跨模態單調對齊與粗粒度多模態語義一致性。CSA通過計算跨模態特徵序列的余弦相似度關聯矩陣,約束符號序列與動作序列的順序一致性;MSC則基於批次數據中的配對與非配對樣本構建多模態三元組,拉近對應的文本-視頻對,推開非對應對,約束符號與動作序列的語義共現程度。在流行的PHOENIX14T基準上的大量實驗表明,LVMCN優於現有最先进方法。

原文摘要 · Abstract (English)

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP suffers huge challenges in linguistics-vision consistency. In this work, we propose a Transformer-based Linguistics-Vision Monotonic Consistent Network (LVMCN) for SLP, which constrains fine-grained cross-modal monotonic alignment and coarse-grained multimodal semantic consistency in language-visual cues through Cross-modal Semantic Aligner (CSA) and Multimodal Semantic Comparator (MSC). In the CSA, we constrain the implicit alignment between corresponding gloss and pose sequences by computing the cosine similarity association matrix between cross-modal feature sequences (i.e., the order consistency of fine-grained sign glosses and actions). As for MSC, we construct multimodal triplets based on paired and unpaired samples in batch data. By pulling closer the corresponding text-visual pairs and pushing apart the non-corresponding text-visual pairs, we constrain the semantic co-occurrence degree between corresponding gloss and pose sequences (i.e., the semantic consistency of coarse-grained textual sentences and sign videos). Extensive experiments on the popular PHOENIX14T benchmark show that the LVMCN outperforms the state-of-the-art.

手语生成跨模态对齐Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。