arXiv:2509.05786cs.MMcs.SD2025-09被引 3

从视频中自动提取音视频文本三模态数据,构建高质量多模态数据集。

Effectively obtaining acoustic, visual and textual data from videos

  • 基于视频自动抽取音频、图像与文本三类数据。
  • 利用图文模型生成描述性文本,确保跨模态语义关联。
  • 公开数据集,适合多模态学习与跨模态分析研究者使用。

机器学习模型的广泛应用加剧了对高质量、大规模多模态数据集的需求。然而,尤其是结合声学、视觉与文本数据的数据集仍十分稀缺。本文提出一种从视频中提取相关音视频文本观测的方法,详细阐述了视频筛选、数据对抽取及利用图像到文本模型生成描述性文本的过程。该方法确保了不同模态间的稳健语义关联,提升了所生成数据集在各类应用中的实用性。同时讨论了实际挑战并提出改进方案以提升数据质量。最终构建的数据集已公开,旨在支持并推动多模态数据分析与机器学习研究。

原文摘要 · Abstract (English)

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains limited. This paper addresses this gap by proposing a method to extract related audio-image-text observations from videos. We detail the process of selecting suitable videos, extracting relevant data pairs, and generating descriptive texts using image-to-text models. Our approach ensures a robust semantic connection between modalities, enhancing the utility of the created datasets for various applications. We also discuss the challenges encountered and propose solutions to improve data quality. The resulting datasets, publicly available, aim to support and advance research in multimodal data analysis and machine learning.

多模态数据视频理解数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。