arXiv:2412.15509cs.CVcs.MM2024-12

微调视觉语言模型,提升视频描述生成的准确性与专业性

PolySmart @ TRECVid 2024 Video Captioning (VTT)

  • 用VTT数据微调LLaVA等模型,增强视频理解与描述能力
  • 微调后模型在多个指标上超越基线,细节更丰富、语义更贴合
  • 适合需要高质量视频描述的应用场景,如媒体内容标注

本文介绍了我们在TRECVid 2024视频到文本(VTT)任务中的方法与结果,探索了LLaVA和LLaVA-NeXT-Video等视觉语言模型在生成视频自然语言描述方面的表现。研究分析了在VTT数据集上微调这些模型对描述准确性、上下文相关性和语言一致性的提升效果。实验表明,微调显著增强了模型生成更详细、领域对齐文本的能力,缩小了通用视觉语言模型与VTT任务特殊需求之间的差距。在多种评估指标下,微调模型均优于基线模型,证明了针对特定任务进行领域微调的重要性。

原文摘要 · Abstract (English)

In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA-NeXT-Video in generating natural language descriptions for video content. We investigate the impact of fine-tuning VLMs on VTT datasets to enhance description accuracy, contextual relevance, and linguistic consistency. Our analysis reveals that fine-tuning substantially improves the model's ability to produce more detailed and domain-aligned text, bridging the gap between generic VLM tasks and the specialized needs of VTT. Experimental results demonstrate that our fine-tuned model outperforms baseline VLMs across various evaluation metrics, underscoring the importance of domain-specific tuning for complex VTT tasks.

视频描述视觉语言模型微调VTT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。