arXiv:2601.15016cs.CV2026-01AAAI被引 6

首个面向互动直播视频的多模态评测基准,解决传统视频评估忽略实时互动的问题。

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

  • 构建包含24项任务的半自动标注流程,融合多模型协作与人工校验
  • 提出VCR模块和两阶段微调,使7B模型性能超越72B开源模型
  • 适合研究直播理解、多模态交互或视频生成的开发者使用

多模态大语言模型(MLLM)推动了通用视频理解的发展。然而现有视频评测基准主要聚焦非互动视频,如电影和录像。为填补这一空白,本文提出首个面向互动直播视频的全模态评测基准LiViBench,涵盖24项任务,突出感知、推理及直播特有挑战。为高效构建数据集,设计标准化的半自动标注流程,在多个阶段引入人机协同;利用多个MLLM构成多智能体系统进行视频全面描述,并采用种子问题驱动方法生成高质量标注。所有互动视频均包含音频、语音与实时评论模态。为增强模型对互动视频的理解,设计定制化的两阶段指令微调,并提出视频到评论检索(VCR)模块以提升模型利用实时评论的能力。基于上述进展,开发了具备直播理解能力的LiVi-LLM-7B模型。实验表明,该模型在性能上超越参数量高达72B的开源模型,缩小与领先专有模型在LiViBench上的差距,并在VideoMME、LongVideoBench、MLVU和VideoEval-Pro等通用视频基准上表现更优。

原文摘要 · Abstract (English)

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for interactive livestream videos, LiViBench. It features a diverse set of 24 tasks, highlighting the perceptual, reasoning, and livestream-specific challenges. To efficiently construct the dataset, we design a standardized semi-automatic annotation workflow that incorporates the human-in-the-loop at multiple stages. The workflow leverages multiple MLLMs to form a multi-agent system for comprehensive video description and uses a seed-question-driven method to construct high-quality annotations. All interactive videos in the benchmark include audio, speech, and real-time comments modalities. To enhance models' understanding of interactive videos, we design tailored two-stage instruction-tuning and propose a Video-to-Comment Retrieval (VCR) module to improve the model's ability to utilize real-time comments. Based on these advancements, we develop LiVi-LLM-7B, an MLLM with enhanced knowledge of interactive livestreams. Experiments show that our model outperforms larger open-source models with up to 72B parameters, narrows the gap with leading proprietary models on LiViBench, and achieves enhanced performance on general video benchmarks, including VideoMME, LongVideoBench, MLVU, and VideoEval-Pro.

直播理解多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。