arXiv:2409.16765cs.CVcs.AI2024-09被引 3

多模态对齐算法提升课件与视频匹配精度,速度快11倍

MaViLS, a Benchmark Dataset for Video-to-Slide Alignment, Assessing Baseline Accuracy with a Multimodal Alignment Algorithm Leveraging Speech, OCR, and Visual Features

  • 融合语音、文字、图像特征,用动态规划找最佳课件序列
  • 平均准确率0.82,远超传统SIFT的0.56,且提速11倍
  • 适合处理视频质量差或讲课风格不一的场景

本文提出一个用于讲座视频与对应课件对齐的基准数据集,并引入一种新型多模态对齐算法,融合语音、文本和视觉特征。该算法利用动态规划确定最优课件序列,在对比SIFT方法(准确率0.56)时达到平均0.82的准确率,且运算速度约快11倍。结果表明,惩罚课件切换可进一步提升准确率;通过光学字符识别(OCR)获取的文本特征对匹配精度贡献最大,其次为图像特征。研究还发现,仅依赖语音转录文本也能提供有效信息,尤其在缺乏OCR数据时。不同课程间的匹配准确率差异反映出视频质量与讲授风格带来的挑战。该多模态算法展现出对部分挑战的鲁棒性,凸显了该方法的潜力。

原文摘要 · Abstract (English)

This paper presents a benchmark dataset for aligning lecture videos with corresponding slides and introduces a novel multimodal algorithm leveraging features from speech, text, and images. It achieves an average accuracy of 0.82 in comparison to SIFT (0.56) while being approximately 11 times faster. Using dynamic programming the algorithm tries to determine the optimal slide sequence. The results show that penalizing slide transitions increases accuracy. Features obtained via optical character recognition (OCR) contribute the most to a high matching accuracy, followed by image features. The findings highlight that audio transcripts alone provide valuable information for alignment and are beneficial if OCR data is lacking. Variations in matching accuracy across different lectures highlight the challenges associated with video quality and lecture style. The novel multimodal algorithm demonstrates robustness to some of these challenges, underscoring the potential of the approach.

视频对齐多模态课件生成OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。