用语义概念评估视频相似性,让模型更像人一样理解视频差异。
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- 基于预定义语义概念计算视频对的可解释相似度。
- 在多领域视频对上验证,发现部分概念更难准确判断。
- 适合研究语言驱动视频理解与可解释性的人看。
两个视频何时算相似?它们可能在动作上相似,但在拍摄地点上完全不同。人类能综合不同方面判断,而现有模型常依赖全局相似分数。具备视频理解能力的大规模多模态模型(LMMs)为利用自然语言进行视频对比带来新可能。我们提出概念式视频相似性估计(ConViS),通过预设关键语义概念,计算视频对间的可解释相似度,实现类人推理,并支持概念条件下的视频检索。为此,我们构建了ConViS-Bench基准,包含跨多个领域的精心标注视频对,每对配有概念级相似度评分及异同文本描述。我们在该基准上评测多个先进模型,发现其与人类判断存在显著差异,表明某些概念更难建模。我们认为ConViS-Bench将推动语言驱动视频理解研究发展。
原文摘要 · Abstract (English)
What does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not been thoroughly studied and presents a challenge for models that often depend on broad global similarity scores. Large Multimodal Models (LMMs) with video understanding capabilities open new opportunities for leveraging natural language in comparative video tasks. We introduce Concept-based Video Similarity estimation (ConViS), a novel task that compares pairs of videos by computing interpretable similarity scores across a predefined set of key semantic concepts. ConViS allows for human-like reasoning about video similarity and enables new applications such as concept-conditioned video retrieval. To support this task, we also introduce ConViS-Bench, a new benchmark comprising carefully annotated video pairs spanning multiple domains. Each pair comes with concept-level similarity scores and textual descriptions of both differences and similarities. Additionally, we benchmark several state-of-the-art models on ConViS, providing insights into their alignment with human judgments. Our results reveal significant performance differences on ConViS, indicating that some concepts present greater challenges for estimating video similarity. We believe that ConViS-Bench will serve as a valuable resource for advancing research in language-driven video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。