arXiv:2502.08279cs.CLcs.AI2025-02ACL被引 18

构建科学演讲视频转文本摘要数据集,助力精准学术内容提炼。

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations

  • 设计VISTA数据集,含1.86万场AI会议演讲与论文摘要
  • 引入规划框架提升摘要结构化与事实一致性
  • 模型与人类仍有显著差距,适合研究多模态摘要

将录制视频转化为简洁准确的文本摘要,是多模态学习中的重要挑战。本文提出VISTA数据集,专为科学领域视频到文本摘要任务设计,包含18,599个经录制的AI会议演讲及其对应的论文摘要。我们对顶尖大模型进行基准测试,并采用基于计划的框架以更好捕捉摘要的结构特征。人工与自动评估均表明,显式规划能有效提升摘要质量与事实一致性。然而,模型性能与人类表现之间仍存在显著差距,凸显了该数据集的挑战性。本研究旨在为未来科学视频转文本摘要的研究铺平道路。

原文摘要 · Abstract (English)

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paired with their corresponding paper abstracts. We benchmark the performance of state-of-the-art large models and apply a plan-based framework to better capture the structured nature of abstracts. Both human and automated evaluations confirm that explicit planning enhances summary quality and factual consistency. However, a considerable gap remains between models and human performance, highlighting the challenges of our dataset. This study aims to pave the way for future research on scientific video-to-text summarization.

视频摘要多模态科学文本数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。