arXiv:2412.14965cs.CVcs.AI2024-12

构建新基准评估视频故事生成,发现大模型表现不佳并提出改进方案

Movie2Story: A framework for understanding videos and telling stories in the form of novel text

  • 设计多模态故事生成基准MSBench,自动构建含辅助信息的长视频数据集
  • 现有多模态大模型在复杂上下文文本生成中表现欠佳,准确率不足预期
  • 提出新模型架构,显著提升长视频故事生成质量,适合视频理解研究者

近年来,大规模模型取得显著进展,伴随大量高质量评估基准的出现。然而,现有基准主要聚焦静态图像中的空间理解,少数扩展至时序任务,仍难以评估长视频与丰富辅助信息下的文本生成能力。为此,我们提出新型基准——多模态故事生成基准(MSBench),用于评估复杂上下文中的文本生成性能。本工作引入创新的自动化数据集生成方法,通过复用现有数据集并自动处理生成新数据集,大幅减少人工成本;同时通过系统化过滤和先进模型校验,确保辅助数据与真实标签的准确性。实验表明,当前多模态大语言模型(MLLMs)在所提评估指标下表现不佳,揭示其能力短板。为此,我们提出一种新型模型架构与方法,有效提升整体处理效果,在基准上实现显著改进。

原文摘要 · Abstract (English)

In recent years, large-scale models have achieved significant advancements, accompanied by the emergence of numerous high-quality benchmarks for evaluating various aspects of their comprehension abilities. However, most existing benchmarks primarily focus on spatial understanding in static image tasks. While some benchmarks extend evaluations to temporal tasks, they fall short in assessing text generation under complex contexts involving long videos and rich auxiliary information. To address this limitation, we propose a novel benchmark: the Multi-modal Story Generation Benchmark (MSBench), designed to evaluate text generation capabilities in scenarios enriched with auxiliary information. Our work introduces an innovative automatic dataset generation method to ensure the availability of accurate auxiliary information. On one hand, we leverage existing datasets and apply automated processes to generate new evaluation datasets, significantly reducing manual efforts. On the other hand, we refine auxiliary data through systematic filtering and utilize state-of-the-art models to ensure the fairness and accuracy of the ground-truth datasets. Our experiments reveal that current Multi-modal Large Language Models (MLLMs) perform suboptimally under the proposed evaluation metrics, highlighting significant gaps in their capabilities. To address these challenges, we propose a novel model architecture and methodology to better handle the overall process, demonstrating improvements on our benchmark.

视频理解故事生成多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。