arXiv:2409.08444cs.CV2024-09被引 4

用大模型统一识别面部动作单元,准确率提升超11%。

Towards Unified Facial Action Unit Recognition Framework by Large Language Models

  • 结合视觉编码器与大语言模型,生成多格式面部动作识别结果。
  • 在BP4D和DISFA数据集上近半数动作单元识别精度领先。
  • 首次实现基于大模型的统一面部动作识别框架,适合情感计算研究者。

面部动作单元(AUs)在情感计算中具有重要意义。本文提出首个基于大语言模型(LLM)的统一面部动作单元识别框架AU-LLaVA,由视觉编码器、线性投影层和预训练大语言模型构成。通过精心设计文本描述并在多个AU数据集上微调,模型可对同一输入图像生成不同格式的识别结果。在BP4D和DISFA数据集上,AU-LLaVA在近一半动作单元的识别中达到最高准确率;相比先前基准,特定动作单元的F1分数提升最高达11.4%。在FEAFA数据集上,方法对全部24个动作单元均取得显著提升。该模型展现出卓越的性能与泛化能力。

原文摘要 · Abstract (English)

Facial Action Units (AUs) are of great significance in the realm of affective computing. In this paper, we propose AU-LLaVA, the first unified AU recognition framework based on the Large Language Model (LLM). AU-LLaVA consists of a visual encoder, a linear projector layer, and a pre-trained LLM. We meticulously craft the text descriptions and fine-tune the model on various AU datasets, allowing it to generate different formats of AU recognition results for the same input image. On the BP4D and DISFA datasets, AU-LLaVA delivers the most accurate recognition results for nearly half of the AUs. Our model achieves improvements of F1-score up to 11.4% in specific AU recognition compared to previous benchmark results. On the FEAFA dataset, our method achieves significant improvements over all 24 AUs compared to previous benchmark results. AU-LLaVA demonstrates exceptional performance and versatility in AU recognition.

面部识别大模型情感计算动作单元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。