arXiv:2606.30001cs.CVcs.GR2026-06中稿 · ECCV

让虚拟人说不同文化的话,做符合文化的动作。

SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset

论文配图:SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
图 1 · 摘自论文原文
  • 用跨说话人文化表征,分离文化与个人动作风格
  • 在106小时数据上提升动作真实感与文化一致性
  • 适合需要跨文化交互的虚拟角色生成应用

近期的伴随语音手势生成方法常忽视文化差异,限制了人机交互效果。此外,文化条件模型很少在说话人独立划分下评估,导致“文化”行为可能混杂说话人个人动作风格。我们提出SICAGE,一种模块化框架,通过说话人无关的文化表征来驱动手势生成。该框架利用音频和文本信息,将每位说话人视为独立域,同时施加跨说话人的不变性,使表征保持文化区分性的同时降低对说话人身份的依赖。由此获得的文化嵌入用于控制多模态生成器,输出符合文化的动作。我们采用对抗学习和Fishr正则化两种领域泛化方法实现该思想,并设计ALaDiT——一个实时扩散模型手势生成器,可高效融入学习到的文化嵌入。为验证方法,我们构建了TED4C-L数据集,包含来自四个文化群体的764位演讲者、总计106小时的多模态数据。实验表明,SICAGE在动作真实度、多样性、节拍同步性、语义相关性及文化一致性方面均有提升。

原文摘要 · Abstract (English)

Recent co-speech gesture generation methods often overlook cultural differences, limiting their effectiveness in human-agent interaction. Moreover, culture-conditioned models are rarely evaluated under speaker-disjoint splits, so apparent "cultural" behavior may be confounded with speaker-specific gesturing style. We introduce SICAGE, a modular framework for culture-aware co-speech gesture generation that conditions motion synthesis models on speaker-independent cultural representations. SICAGE learns these representations from audio and text by treating each speaker as a separate domain while imposing invariance across speakers. This encourages representations to remain culture-discriminative while reducing dependence on speaker identity. The resulting cultural embeddings condition a multimodal generator to produce culturally appropriate gestures. We instantiate this idea with two domain generalization approaches: adversarial learning and Fishr regularization. We further introduce ALaDiT, a real-time diffusion-based gesture generator designed to efficiently incorporate the learned cultural embeddings. To validate our method, we built TED4C-L, a 106-hour multimodal dataset of 764 TED speakers from four cultural groups. Experiments show that SICAGE improves motion realism, diversity, beat synchronization, semantic relevance, and cultural consistency.

手势生成文化差异扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。