arXiv:2503.17069cs.CVcs.AI2025-03ICCV被引 11

仅用一段视频就能让模型认出特定人物并回答相关问题。

PVChat: Personalized Video Chat with One-Shot Learning

  • 通过单视频学习构建个性化视觉语言模型,支持人物识别与问答。
  • 在医疗、影视等场景中,单视频训练后准确率显著优于现有模型。
  • 创新设计路由机制和增强损失函数,提升模型对个体特征的敏感度。

视频大语言模型(ViLLMs)在通用视频理解方面表现优异,如识别说话、进食等行为,但在涉及身份认知的理解任务中表现不佳,例如‘威尔逊正在接受化疗’或‘汤姆正与莎拉讨论’,限制了其在智慧医疗和智能家居中的应用。为此,我们提出首个基于单次学习的个性化视频大模型框架PVChat,实现从单段视频中对每个主体进行身份感知的问答。方法上,在合成增强的视频-问答数据集上优化改进的混合头(MoH)ViLLM,采用渐进式图像到视频学习策略。引入自动化数据增强流程,生成保持身份一致的正样本,并从现有视频语料库中检索难例负样本,构建涵盖存在性、外观、动作、位置四类问题的多样化训练数据集。为强化个体特征学习,提出基于ReLU路由的MoH注意力机制及两项新目标:平滑邻近正则化(通过指数距离缩放实现渐进学习)与头激活增强(均衡注意力分配)。最终采用两阶段训练策略,从图像预训练过渡到视频微调,实现从静态属性到动态表征的逐步学习。在覆盖医疗场景、电视剧、动漫和真实影像的多个数据集上评估,结果表明,仅需单视频输入,PVChat即在个性化特征理解上超越当前最优ViLLMs。

原文摘要 · Abstract (English)

Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthcare and smart home environments. To address this limitation, we propose a one-shot learning framework PVChat, the first personalized ViLLM that enables subject-aware question answering (QA) from a single video for each subject. Our approach optimizes a Mixture-of-Heads (MoH) enhanced ViLLM on a synthetically augmented video-QA dataset, leveraging a progressive image-to-video learning strategy. Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. To enhance subject-specific learning, we propose a ReLU Routing MoH attention mechanism, alongside two novel objectives: (1) Smooth Proximity Regularization for progressive learning through exponential distance scaling and (2) Head Activation Enhancement for balanced attention routing. Finally, we adopt a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations. We evaluate PVChat on diverse datasets covering medical scenarios, TV series, anime, and real-world footage, demonstrating its superiority in personalized feature understanding after learning from a single video, compared to state-of-the-art ViLLMs.

视频理解个性化单次学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。