arXiv:2510.11129cs.CVcs.AI2025-10被引 2

让AI持续理解3小时视频,还能从记忆中学习新知识。

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

  • 用测试时训练构建可更新的长期记忆,替代传统压缩方法。
  • 在3小时视频上准确率比非流式模型高15%,长时任务提升7%。
  • 适合需要持续学习的智能体、视频分析等场景。

长时序流式视频理解对未来AI智能体至关重要,但受限于低效的长期记忆机制。我们提出video-SALMONN S,一个增强记忆的流式音视频大语言模型,可在1 FPS、360p分辨率下处理超过3小时的视频,在相同内存预算下超越强基准非流式模型。除令牌合并与降采样外,video-SALMONN S首次将测试时训练(TTT)作为流式记忆机制,持续将短期多模态表征转化为嵌入模型参数的长期记忆。为提升长程依赖建模与记忆容量,我们提出:(i) 增加长跨度预测目标的TTT_MEM层;(ii) 两阶段训练策略;(iii) 模态感知记忆读取器。我们还构建了模拟智能体行为的Episodic Learning from Video Memory(ELViM)基准,要求模型在数小时后仍能从观看过的视频中学习。video-SALMONN S在长视频基准上始终优于流式与非流式基线3-7%,在ELViM上相较强非流式模型绝对准确率提升15%,展现强大视频记忆学习能力。

原文摘要 · Abstract (English)

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language model that processes over 3-hour videos at 1 FPS and 360p resolution, outperforming strong non-streaming models under the same memory budget. In addition to token merging or downsampling, video-SALMONN S is the first to employ test-time training (TTT) as a streaming memory mechanism for video understanding. TTT continuously transforms short-term multimodal representations into long-term memory embedded in model parameters. To improve long-range dependency modeling and memory capacity, we propose (i) a TTT_MEM layer with an additional long-span prediction objective, (ii) a two-stage training scheme, and (iii) a modality-aware memory reader. We further introduce the Episodic Learning from Video Memory (ELViM) benchmark, simulating agent-like scenarios where models must learn from videos observed hours earlier. video-SALMONN S consistently outperforms both streaming and non-streaming baselines by 3-7% on long video benchmarks. Notably, video-SALMONN S achieves a 15% absolute accuracy improvement over strong non-streaming models on ELViM, demonstrating strong learning abilities from video memory.

视频理解长时记忆流式处理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。