arXiv:2509.26225cs.CVcs.AI2025-09中稿 · version

用大模型生成视频摘要的合理文字解释,验证其可信度。

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

  • 用LLaVA-OneVision生成视觉解释的文字描述
  • 通过语义重叠衡量解释与摘要的一致性,评估可信度
  • 在SumMe和TVSum数据集上验证方法有效性

本文针对视频摘要结果生成合理文本解释开展实验研究。在现有多粒度解释框架基础上,集成当前最优的大规模多模态模型LLaVA-OneVision,通过提示(prompting)生成视觉解释的自然语言描述。研究聚焦可解释AI中关键特性——解释的可信度,即是否符合人类推理与预期。基于扩展框架,提出一种评估方法:利用SBERT与SimCSE两种句向量方法,量化视觉解释的文字描述与对应视频摘要文字描述之间的语义重叠。在当前最优方法CA-SUM及两个数据集SumMe、TVSum上进行实验,检验更忠实的解释是否也更具可信度,并识别生成可信文本解释的最佳方案。

原文摘要 · Abstract (English)

In this paper, we present our experimental study on generating plausible textual explanations for the outcomes of video summarization. For the needs of this study, we extend an existing framework for multigranular explanation of video summarization by integrating a SOTA Large Multimodal Model (LLaVA-OneVision) and prompting it to produce natural language descriptions of the obtained visual explanations. Following, we focus on one of the most desired characteristics for explainable AI, the plausibility of the obtained explanations that relates with their alignment with the humans' reasoning and expectations. Using the extended framework, we propose an approach for evaluating the plausibility of visual explanations by quantifying the semantic overlap between their textual descriptions and the textual descriptions of the corresponding video summaries, with the help of two methods for creating sentence embeddings (SBERT, SimCSE). Based on the extended framework and the proposed plausibility evaluation approach, we conduct an experimental study using a SOTA method (CA-SUM) and two datasets (SumMe, TVSum) for video summarization, to examine whether the more faithful explanations are also the more plausible ones, and identify the most appropriate approach for generating plausible textual explanations for video summarization.

视频摘要可解释AI大模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。