arXiv:2603.26553cs.CV2026-03被引 4

用对比流匹配生成更自然的伴随言语手势,避免机械重复动作。

SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

  • 引入负样本对比学习,让手势更符合语义而非仅模仿节奏。
  • 在两个数据集上优于现有方法,用户评估认可手势自然度。
  • 将语音、文本与全身动作统一建模,保持多模态一致性。

尽管伴随言语手势生成领域已取得显著进展,但生成整体性且语义一致的手势仍具挑战。现有方法依赖外部语义检索,受限于预设语言规则,泛化能力弱。基于流匹配的方法虽表现良好,但仅使用语义一致样本训练,缺乏负例,导致学习到的是重复性节奏动作,而非象征性或隐喻性等稀疏手势。此外,多数方法孤立建模身体部位,难以保证跨模态一致性。本文提出一种基于对比流匹配的伴随言语手势生成模型,利用语义不匹配的音视频条件作为负样本,使速度场既能沿正确轨迹运动,又能排斥语义不符的轨迹。通过余弦与对比目标,将文本、音频和整体动作嵌入复合潜在空间,确保跨模态一致性。大量实验与用户研究显示,该方法在BEAT2和SHOW两个数据集上均超越现有最优模型。

原文摘要 · Abstract (English)

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their generalisation capability due to dependency on predefined linguistic rules. Flow-matching-based methods produce promising results; however, the network is optimised using only semantically congruent samples without exposure to negative examples, leading to learning rhythmic gestures rather than sparse motion, such as iconic and metaphoric gestures. Furthermore, by modelling body parts in isolation, the majority of methods fail to maintain crossmodal consistency. We introduce a Contrastive Flow Matching-based co-speech gesture generation model that uses mismatched audio-text conditions as negatives, training the velocity field to follow the correct motion trajectory while repelling semantically incongruent trajectories. Our model ensures cross-modal coherence by embedding text, audio, and holistic motion into a composite latent space via cosine and contrastive objectives. Extensive experiments and a user study demonstrate that our proposed approach outperforms state-of-the-art methods on two datasets, BEAT2 and SHOW.

手势生成流匹配多模态对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。