arXiv:2503.09158cs.CV2025-03被引 5

针对人脸视频理解中线索丢失问题,提出分层提示引导的视觉推理模型

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO

  • 分层提示引导提取面部动态特征,聚焦问题相关线索
  • 在17万问答对上实现零样本任务性能超越现有模型
  • 适用于需要精细面部细节理解的任务,如情感分析、微表情识别

现有视频大语言模型主要依赖无提示感知的视觉编码器,提取无关紧要的人脸表征,导致关键线索丢失。为此,我们提出FaVChat,首个面向细微视觉与动态面部线索推理的视频大语言模型。FaVChat引入分层、提示引导的视觉特征提取框架,在三个互补层次强调与问题相关的信息,并动态融合注入大语言模型,提升面部细节推理精度。为应对数据稀缺下的学习效率问题,提出数据高效GRPO(Data-Efficient GRPO)强化学习策略,通过逐样本效用估计迭代识别高价值样本,最大化每条数据贡献。构建大规模基准数据集FaVChat 170K,包含约6万高质量人脸视频和17万问答对,聚焦细粒度面部细节。大量实验表明,包括在四个面部理解任务上的零样本评估,FaVChat持续优于现有VLLMs。

原文摘要 · Abstract (English)

Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried information, leading to the loss of task critical cues. To address this challenge, we propose FaVChat, the first VLLM designed for reasoning over subtle visual and dynamic facial cues. FaVChat introduces a hierarchical, prompt guided visual feature extraction framework that emphasizes question relevant information at three complementary levels. These multi level features are dynamically fused and injected into the LLM, enabling more accurate facial details reasoning To further improve learning efficiency under data scarcity, we propose Data Efficient GRPO, a reinforcement learning strategy that iteratively identifies high utility samples and maximizes the contribution of each instance via per instance utility estimation, substantially enhancing performance gains under limited supervision. We construct a large scale benchmark dataset FaVChat 170K, comprising approximately 60K high quality facial videos and 170K question answer pairs focusing on fine grained facial details. Extensive experiments, including zero shot evaluations on four facial understanding tasks, demonstrate that FaVChat consistently outperforms existing VLLMs.

视频理解人脸分析提示引导强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。