构建新数据集与模型,提升视频中人脸表情理解能力
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
- 构建5033段高质量视频表情数据集,含70万+标注令牌
- 提出FaceTrack-MM模型,用少量令牌精准追踪主角色面部表情
- 设计新评估指标与FEC-Bench基准,兼顾内容与时间一致性
人脸表情描述在多个领域已有广泛应用。近年来,视频多模态大语言模型(MLLM)在通用视频理解任务中展现出潜力,但在视频中描述人脸表情仍面临两大挑战:(1)缺乏充足的数据集与评测基准;(2)视频MLLM的视觉令牌容量有限。为此,本文引入一个面向动态人脸表情描述的指令跟随数据集,包含5,033段人工标注的高质量视频片段,总标注令牌超过70万。该数据集旨在提升视频MLLM对细微面部变化的感知能力。此外,我们提出FaceTrack-MM模型,仅用少量令牌即可编码主要人物面部特征,在复杂多人场景中仍能有效追踪并聚焦主体面部表情。同时,我们设计了一种结合事件抽取、关系分类与最长公共子序列(LCS)算法的新评估指标,用于衡量生成文本的内容一致性和时序连贯性。此外,我们构建了FEC-Bench基准,用于评估现有视频MLLM在此任务上的表现。所有数据与源代码将公开共享。
原文摘要 · Abstract (English)
Facial expression captioning has found widespread application across various domains. Recently, the emergence of video Multimodal Large Language Models (MLLMs) has shown promise in general video understanding tasks. However, describing facial expressions within videos poses two major challenges for these models: (1) the lack of adequate datasets and benchmarks, and (2) the limited visual token capacity of video MLLMs. To address these issues, this paper introduces a new instruction-following dataset tailored for dynamic facial expression caption. The dataset comprises 5,033 high-quality video clips annotated manually, containing over 700,000 tokens. Its purpose is to improve the capability of video MLLMs to discern subtle facial nuances. Furthermore, we propose FaceTrack-MM, which leverages a limited number of tokens to encode the main character's face. This model demonstrates superior performance in tracking faces and focusing on the facial expressions of the main characters, even in intricate multi-person scenarios. Additionally, we introduce a novel evaluation metric combining event extraction, relation classification, and the longest common subsequence (LCS) algorithm to assess the content consistency and temporal sequence consistency of generated text. Moreover, we present FEC-Bench, a benchmark designed to assess the performance of existing video MLLMs in this specific task. All data and source code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。