用大模型分阶段评估视频段落与用户主题的关联度。
LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content
- 分两步评分:先看单段内容,再结合前后文上下文。
- 输出带分数标注的视频和分析日志。
- 适合需要精准追踪主题变化的多媒体分析场景。
本文提出学习型用户重要性追踪框架LUST,用于分析视频内容并量化其片段与用户提供的文本主题之间的语义相关性。该框架采用多模态分析流程,融合视频帧的视觉信息与语音识别(ASR)提取的音频文本信息。核心创新在于基于大语言模型(LLM)的分层两级相关性评分机制:第一阶段计算直接相关性得分$S_{d,i}$,评估单个片段在视觉与听觉层面与主题的即时匹配度;第二阶段计算上下文相关性得分$S_{c,i}$,通过整合前序片段的主题得分序列,捕捉叙事的动态演变过程。该框架旨在提供一种具有时间感知能力的、细致的用户定义重要性度量,输出带有可视化相关性分数的标注视频及完整分析日志。
原文摘要 · Abstract (English)
This paper introduces the Learned User Significance Tracker (LUST), a framework designed to analyze video content and quantify the thematic relevance of its segments in relation to a user-provided textual description of significance. LUST leverages a multi-modal analytical pipeline, integrating visual cues from video frames with textual information extracted via Automatic Speech Recognition (ASR) from the audio track. The core innovation lies in a hierarchical, two-stage relevance scoring mechanism employing Large Language Models (LLMs). An initial "direct relevance" score, $S_{d,i}$, assesses individual segments based on immediate visual and auditory content against the theme. This is followed by a "contextual relevance" score, $S_{c,i}$, that refines the assessment by incorporating the temporal progression of preceding thematic scores, allowing the model to understand evolving narratives. The LUST framework aims to provide a nuanced, temporally-aware measure of user-defined significance, outputting an annotated video with visualized relevance scores and comprehensive analytical logs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。