arXiv:2501.07888cs.CVcs.AI2025-01被引 80

Tarsier2提升视频理解能力,细节描述更准,通用性更强。

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

  • 扩大训练数据至4000万对视频文本,增强多样性与细节捕捉能力。
  • 在微调中进行精细时间对齐,提升时序理解精度。
  • 用模型生成偏好数据并优化,让视频描述更贴合人类判断。

我们提出Tarsier2,一种先进的大型视觉语言模型(LVLM),可生成详细准确的视频描述,并具备出色的通用视频理解能力。通过三大升级实现显著进步:(1)将预训练数据规模从1100万扩大至4000万视频-文本对,提升数量与多样性;(2)在监督微调中实施细粒度时间对齐;(3)使用模型生成偏好数据并采用DPO训练进行优化。大量实验表明,Tarsier2-7B在详细视频描述任务中持续优于GPT-4o和Gemini 1.5 Pro等领先商用模型。在DREAM-1K基准上,其F1得分比GPT-4o高2.8%,比Gemini 1.5 Pro高5.8%。人工对比评估中,相比GPT-4o提升8.6%,比Gemini 1.5 Pro高24.9%。Tarsier2-7B在15个公开基准上取得新最优表现,涵盖视频问答、视频定位、幻觉测试及具身问答等任务,展现其作为强大通用视觉语言模型的多功能性。

原文摘要 · Abstract (English)

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tarsier2 achieves significant advancements through three key upgrades: (1) Scaling pre-training data from 11M to 40M video-text pairs, enriching both volume and diversity; (2) Performing fine-grained temporal alignment during supervised fine-tuning; (3) Using model-based sampling to automatically construct preference data and applying DPO training for optimization. Extensive experiments show that Tarsier2-7B consistently outperforms leading proprietary models, including GPT-4o and Gemini 1.5 Pro, in detailed video description tasks. On the DREAM-1K benchmark, Tarsier2-7B improves F1 by 2.8% over GPT-4o and 5.8% over Gemini-1.5-Pro. In human side-by-side evaluations, Tarsier2-7B shows a +8.6% performance advantage over GPT-4o and +24.9% over Gemini-1.5-Pro. Tarsier2-7B also sets new state-of-the-art results across 15 public benchmarks, spanning tasks such as video question-answering, video grounding, hallucination test, and embodied question-answering, demonstrating its versatility as a robust generalist vision-language model.

视频理解视觉语言模型大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。