arXiv:2410.03051cs.CV2024-10ICLR被引 150

提出高效视频细粒度描述模型AuroraCap及新基准VDC,性能超越GPT-4V。

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

  • 基于大模型设计简洁架构,用标记合并减少视觉令牌数
  • 在Flickr30k上达CIDEr 88.9,优于GPT-4V和Gemini-1.5 Pro
  • 构建超千条结构化描述的新基准,支持高质量评估

视频细粒度描述旨在生成全面连贯的视频文本描述,对视频理解与生成均有裨益。本文提出基于大模态模型的AuroraCap,采用最简架构,不引入额外时序参数。为应对长视频序列带来的计算开销,采用标记合并策略,显著减少输入视觉令牌数量,意外发现性能损失极小。AuroraCap在多个视频与图像描述基准上表现优异,例如在Flickr30k上达到CIDEr 88.9,超越GPT-4V(55.3)和Gemini-1.5 Pro(82.2)。然而,现有视频描述基准仅含简短描述(数十词),限制了该领域研究。为此,我们构建了包含超过一千条精心标注结构化描述的VDC新基准。同时提出新的LLM辅助评估指标VDCscore,采用分治策略将长描述评估转化为多个短问答对。结合人工埃洛排名实验,验证该基准更贴近人类对视频细粒度描述质量的判断。

原文摘要 · Abstract (English)

Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality.

视频生成多模态评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。