构建首个面向UI操作视频的多模态摘要数据集,助力精准教学指导。
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
- 构建2413段UI操作视频数据集,含视频分段与文本摘要标注
- 现有方法在生成可执行步骤指令上表现差,准确率不足60%
- 适合研究教学视频生成、人机交互及自动化教程系统者参考
我们研究面向教学视频的多模态摘要任务,旨在以文本指令和关键帧形式为用户提供高效的学习方式。现有基准主要聚焦通用语义级视频摘要,难以支持逐步可执行的操作指令和配套图示,而这正是教学视频的核心需求。为此,我们提出一个全新的面向用户界面(UI)教学视频摘要的基准,构建了包含2,413段视频、总时长超167小时的数据集。所有视频均经人工标注视频分割、文本摘要与视频摘要,支持对简洁且可执行摘要的全面评估。我们在所建MS4UI数据集上开展大量实验,结果表明当前最优多模态摘要方法在UI视频摘要任务上表现不佳,凸显了针对此类视频设计新方法的重要性。
原文摘要 · Abstract (English)
We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic semantic-level video summarization, and are not suitable for providing step-by-step executable instructions and illustrations, both of which are crucial for instructional videos. We propose a novel benchmark for user interface (UI) instructional video summarization to fill the gap. We collect a dataset of 2,413 UI instructional videos, which spans over 167 hours. These videos are manually annotated for video segmentation, text summarization, and video summarization, which enable the comprehensive evaluations for concise and executable video summarization. We conduct extensive experiments on our collected MS4UI dataset, which suggest that state-of-the-art multi-modal summarization methods struggle on UI video summarization, and highlight the importance of new methods for UI instructional video summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。