arXiv:2508.12522cs.CV2025-08被引 5

针对个体差异,用多模态数据精准适配表情识别模型

MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training

  • 多模态协同选择源个体,分阶段优化目标个体特征
  • 在三个数据集上准确率超越现有方法5.2%以上
  • 适合医疗情感评估等需个性化建模的场景

个性化表情识别(ER)需将模型适配至个体数据以应对人际差异。现有多源域适应(MSDA)方法常忽略多模态信息或混合所有源为单一域,限制了个体多样性且未显式捕捉个体特征。为此,提出MuSACo:基于协同训练的多模态个体特定选择与自适应方法。该方法利用多模态间互补信息及多源域,实现个体化适配。通过主导模态生成伪标签进行类别感知学习,并结合类别无关损失从低置信度目标样本中学习;同时对各模态源特征进行对齐,仅融合高置信度目标特征。在BioVid、StressID和BAH三个挑战性多模态ER数据集上的实验表明,MuSACo显著优于统一域适应(UDA)及先进MSDA方法。

原文摘要 · Abstract (English)

Personalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable interpersonal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods, where each domain corresponds to a specific subject to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multimodal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets: BioVid, StressID, and BAH show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods.

表情识别多模态个性化建模协同训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。