arXiv:2604.25255cs.CV2026-04

让人脸表情更自然,说话口型不变。

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation

论文配图:Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation
图 1 · 摘自论文原文
  • 用个性化提示词增强视觉与语义关联
  • 通过特征差分缩小跨模态差异
  • 可无缝接入现有表情生成模型

语音保持的人脸表情操控(SPFEM)旨在不改变原语音对应嘴部动作的前提下提升表达力。该任务主要受限于成对数据稀缺——即同一人说同样内容但表情不同的视频帧难以获取,导致情感操控缺乏直接监督。尽管当前视觉-语言模型(VLM)能提取对齐的视觉与语义特征,具备潜在监督价值,但其直接应用仍受限。为此,本文提出个性化跨模态情感关联学习(PCMECL)算法,通过两大改进优化VLM监督:首先,传统VLM对每种情绪使用单一通用提示词,无法捕捉个体表达差异;PCMECL通过引入个体视觉信息来学习个性化提示词,实现更精细的视觉-语义关联。其次,即便经过个性化处理,视觉与语义特征分布仍存在固有偏差;为此,PCMECL采用特征差分策略,将视觉特征变化与语义特征变化进行匹配,从而提供更精确的监督信号。作为即插即用模块,PCMECL可无缝集成至现有SPFEM模型中。在多个数据集上的大量实验验证了该方法的优越性。

原文摘要 · Abstract (English)

Speech-preserving facial expression manipulation (SPFEM) aims to enhance human expressiveness without altering mouth movements tied to the original speech. A primary challenge in this domain is the scarcity of paired data, namely aligned frames of the same individual with identical speech but different expressions, which impedes direct supervision for emotional manipulation. While current Visual-Language Models (VLMs) can extract aligned visual and semantic features, making them a promising source of supervision, their direct application is limited. To this end, we propose a Personalized Cross-Modal Emotional Correlation Learning (PCMECL) algorithm that refines VLM-based supervision through two major improvements. First, standard VLMs rely on a single generic prompt for each emotion, failing to capture expressive variations among individuals. PCMECL addresses this limitation by conditioning on individual visual information to learn personalized prompts, thereby establishing more fine-grained visual-semantic correlations. Second, even with personalization, inherent discrepancies persist between the visual and semantic feature distributions. To bridge this modality gap, PCMECL employs feature differencing to correlate the modalities, providing more precisely aligned supervision by matching the change in visual features to the change in semantic features. As a plug-and-play module, PCMECL can be seamlessly integrated into existing SPFEM models. Extensive experiments across various datasets demonstrate the superior efficacy of our algorithm.

表情生成跨模态学习个性化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。