AffectVLM通过增强文本提示和多视角融合,提升3D/4D表情识别性能。
Contrastive Language-Image Learning with Augmented Textual Prompts for 3D/4D FER Using Vision-Language Model
- 联合表示学习+梯度友好损失,加速视觉特征收敛。
- 使用增强文本提示与混合视角增强,提升模型语义理解能力。
- 支持实时交互推理,适合需要高精度表情分析的应用场景。
本文提出AffectVLM,一种用于3D/4D面部表情识别的视觉语言模型,通过多视角融合实现语义丰富、视觉全面的表情理解。为有效捕捉视觉特征,设计了联合表示学习框架与新型梯度友好损失函数,加速模型向最优特征表示收敛。引入增强文本提示以提升模型语言能力,并采用混合视角增强扩展视觉数据集。同时开发基于Streamlit的实时交互推理应用,支持分布式学习。大量实验验证AffectVLM在多个基准上的优异表现。
原文摘要 · Abstract (English)
In this paper, we introduce AffectVLM, a vision-language model designed to integrate multiviews for a semantically rich and visually comprehensive understanding of facial emotions from 3D/4D data. To effectively capture visual features, we propose a joint representation learning framework paired with a novel gradient-friendly loss function that accelerates model convergence towards optimal feature representation. Additionally, we introduce augmented textual prompts to enhance the model's linguistic capabilities and employ mixed view augmentation to expand the visual dataset. We also develop a Streamlit app for a real-time interactive inference and enable the model for distributed learning. Extensive experiments validate the superior performance of AffectVLM across multiple benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。