arXiv:2510.12299cs.IR2025-10

对比视频问答中不同表示方法,发现视觉帧最准但耗资源,字幕轻量高效

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

  • 对比单模态与多模态输入,系统评估视觉、字幕、音频等表示方式
  • 视觉帧提升准确率但增加显存和推理延迟,字幕在长视频中表现优异
  • 为资源受限的视频问答系统设计提供实证指导

多模态大语言模型在联合处理视觉、文本和音频信息方面,显著提升了视频问答(VideoQA)性能。然而,目前尚不明确哪种视频表示最适合多模态大语言模型,以及不同模态如何在任务准确率与计算效率间权衡。本文针对该问题,对多模态大语言模型在视频问答中的视频表示方法进行了全面的实证研究。我们在两个常用基准数据集 VideoMME 和 LongVideoBench 上,系统评估了仅文本、字幕、视觉帧、音频信号等单模态输入,以及其组合形式的多模态输入。结果表明,视觉帧可显著提升准确率,但带来巨大的显存占用和推理延迟;而字幕则提供一种轻量但有效的替代方案,尤其适用于长视频场景。研究揭示了效果与效率之间的明确权衡,为设计资源敏感型多模态大语言模型视频问答系统提供了实用洞见。

原文摘要 · Abstract (English)

Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. However, it remains unclear which video representations are most effective for MLLMs, and how different modalities balance task accuracy against computational efficiency. In this work, we present a comprehensive empirical study of video representation methods for VideoQA with MLLMs. We systematically evaluate single modality inputs question only, subtitles, visual frames, and audio signals as well as multimodal combinations, on two widely used benchmarks: VideoMME and LongVideoBench. Our results show that visual frames substantially enhance accuracy but impose heavy costs in GPU memory and inference latency, while subtitles provide a lightweight yet effective alternative, particularly for long videos. These findings highlight clear trade-offs between effectiveness and efficiency and provide practical insights for designing resource-aware MLLM-based VideoQA systems.

视频问答多模态表示学习效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。