提出高效视频定位框架,用少内存实现高精度内容识别。
See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval
- 用查询引导的标题编码语义,对齐用户意图
- 通过重要性调制突出相关片段,减少冗余
- 自适应压缩帧,兼顾精度与内存效率
多模态大语言模型在图像识别与推理上取得进展,但视频任务受限于密集帧处理带来的内存压力。现有视频时段检索(VMR)方法依赖稀疏采样,可能丢失信息,尤其在长视频中。我们提出SMORE(See MORE, store less)框架,在保持高信息分辨率的同时提升内存效率。SMORE(1)利用查询引导的标题编码语义,与用户意图对齐;(2)采用查询感知的重要性调制,突出相关片段;(3)自适应压缩帧,保留关键内容并减少冗余。该方法可在不超内存预算的前提下实现高效视频理解。实验表明,SMORE在QVHighlights、Charades-STA和ActivityNet-Captions三个基准上达到当前最优性能。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have improved image recognition and reasoning, but video-related tasks remain challenging due to memory constraints from dense frame processing. Existing Video Moment Retrieval (VMR) methodologies rely on sparse frame sampling, risking potential information loss, especially in lengthy videos. We propose SMORE (See MORE, store less), a framework that enhances memory efficiency while maintaining high information resolution. SMORE (1) uses query-guided captions to encode semantics aligned with user intent, (2) applies query-aware importance modulation to highlight relevant segments, and (3) adaptively compresses frames to preserve key content while reducing redundancy. This enables efficient video understanding without exceeding memory budgets. Experimental validation reveals that SMORE achieves state-of-the-art performance on QVHighlights, Charades-STA, and ActivityNet-Captions benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。