用视觉语言模型总结多模态演示文稿,发现结构化输入效果最佳。
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
- 用分块幻灯片与字幕组合的结构化输入提升摘要质量。
- 从视频中提取幻灯片比直接用原始视频更高效准确。
- 适合需要高效处理长时多模态文档的研究者参考。
视觉语言模型(VLMs)可处理文本、图像、图文交错或长达一小时的视频等多种格式。本文对使用不同输入表示的多模态演示文稿自动摘要进行了细粒度的定量与定性分析。实验表明,在不同输入长度预算下,采用幻灯片提取自视频流作为输入,相较于原始视频更具优势;而图文交错的结构化表示(幻灯片+转录文本)能实现最佳性能。研究还探讨了多模态交互的本质,并提出改进VLM理解此类文档能力的建议。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) can process visual and textual information in multiple formats: texts, images, interleaved texts and images, or even hour-long videos. In this work, we conduct fine-grained quantitative and qualitative analyses of automatic summarization of multimodal presentations using VLMs with various representations as input. From these experiments, we suggest cost-effective strategies for generating summaries from text-heavy multimodal documents under different input-length budgets using VLMs. We show that slides extracted from the video stream can be beneficially used as input against the raw video, and that a structured representation from interleaved slides and transcript provides the best performance. Finally, we reflect and comment on the nature of cross-modal interactions in multimodal presentations and share suggestions to improve the capabilities of VLMs to understand documents of this nature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。