arXiv:2510.06973cs.CV2025-10被引 1

利用大模型内在能力,显著提升长视频字幕中人物身份匹配准确率。

Addressing the ID-Matching Challenge in Long Video Captioning

  • 基于大模型视觉先验,设计新方法挖掘其自身身份识别能力
  • 在GPT-4o上使身份匹配精确率从50%提至90%,召回率从15%提至80%
  • 适用于需连续追踪多人的长视频理解场景

生成长而复杂的视频字幕具有重要意义,尤其在文本到视频生成和多模态理解领域。其中关键挑战是准确识别不同帧中同一人物的身份,即身份匹配(ID-Matching)问题。以往工作对此关注较少,且多依赖点对点匹配,泛化性差。本文提出新思路:利用大视觉语言模型(LVLM)的强先验能力,挖掘其内在的身份匹配潜力。我们构建了首个评估视频字幕身份匹配能力的新基准,通过分析GPT-4o发现,提升图像信息利用和增加个体描述量可显著改善匹配性能。基于此,提出新方法RICE(Recognizing Identities for Captioning Effectively)。大量实验表明,该方法在字幕质量和身份匹配上均优于基线。尤其在GPT-4o上,精度从50%提升至90%,召回率从15%提升至80%,实现长视频中人物的持续跟踪。

原文摘要 · Abstract (English)

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal understanding. One key challenge in long video captioning is accurately recognizing the same individuals who appear in different frames, which we refer to as the ID-Matching problem. Few prior works have focused on this important issue. Those that have, usually suffer from limited generalization and depend on point-wise matching, which limits their overall effectiveness. In this paper, unlike previous approaches, we build upon LVLMs to leverage their powerful priors. We aim to unlock the inherent ID-Matching capabilities within LVLMs themselves to enhance the ID-Matching performance of captions. Specifically, we first introduce a new benchmark for assessing the ID-Matching capabilities of video captions. Using this benchmark, we investigate LVLMs containing GPT-4o, revealing key insights that the performance of ID-Matching can be improved through two methods: 1) enhancing the usage of image information and 2) increasing the quantity of information of individual descriptions. Based on these insights, we propose a novel video captioning method called Recognizing Identities for Captioning Effectively (RICE). Extensive experiments including assessments of caption quality and ID-Matching performance, demonstrate the superiority of our approach. Notably, when implemented on GPT-4o, our RICE improves the precision of ID-Matching from 50% to 90% and improves the recall of ID-Matching from 15% to 80% compared to baseline. RICE makes it possible to continuously track different individuals in the captions of long videos.

视频字幕身份匹配大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。