从人类视角出发,构建视频理解的观看-记忆-推理框架。
Watch, Remember, Reason: Human-View Video Understanding with MLLMs

- 以观看、记忆、推理三能力构建视频理解新范式
- 揭示长视频处理中感知、记忆与推理的核心挑战
- 适合研究多模态大模型视频理解的学者参考
视频理解正因多模态大语言模型(MLLMs)快速演进,研究从短片段转向长时序、多模态和知识密集型场景。此类场景要求模型应对稀疏证据、长程依赖、多模态对齐及有限算力下的可靠推理。本文提出人类视角下的基于LLM的视频理解框架,围绕观看、记忆、推理三项核心能力展开。该框架不将视频任务视为孤立基准,而是统一分析模型如何获取证据、保持上下文并生成有依据输出。我们提出以感知表征、记忆状态、推理轨迹和最终预测刻画视频理解系统,并识别出时空感知、高效长视频处理、记忆建模、流式理解与忠实推理等关键挑战。代表性方法按其在系统中的角色归类:观看涵盖细粒度、全面性、音视频联合与高效感知;记忆包括离线与流式记忆;推理涵盖纯文本推理与视频思维。进一步考察了第一人称、体育、教学、医疗与叙事视频等应用场景,覆盖训练数据集、评估基准在任务类型、监督形式、模态与能力维度上的多样性。最后,提出可扩展、内存感知、证据驱动的视频智能未来方向。相关工作持续更新于 https://github.com/marinero4972/Awesome-HumanView-VideoUnderstanding。
原文摘要 · Abstract (English)
Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios. These scenarios require models to handle sparse evidence, long-range dependencies, multimodal alignment, and reliable inference under limited computational budgets. This work presents a human-view perspective on LLM-based video understanding, organized around three functional abilities: watching, remembering, and reasoning. Rather than treating video tasks as isolated benchmarks, this view provides a unified structure for analyzing how video MLLMs acquire evidence, preserve context, and produce grounded outputs. We introduce a formulation that characterizes video understanding systems by their perceptual representations, memory states, reasoning traces, and final predictions. Based on this formulation, we identify challenges in spatio-temporal perception, efficient long-video processing, memory modeling, streaming understanding, and faithful reasoning. Representative methods are organized by their roles in video MLLM systems. Watching covers fine-grained, comprehensive, audio-visual, and efficient perception. Remembering includes offline and streaming memory, while reasoning covers text-only reasoning and thinking with videos. We further examine application domains such as egocentric, sports, instructional, medical, and narrative videos, and cover training datasets and evaluation benchmarks across task types, supervision formats, modalities, and capability dimensions. Finally, we outline open problems and future directions for scalable, memory-aware, and evidence-grounded video intelligence. Related works will be continuously traced at https://github.com/marinero4972/Awesome-HumanView-VideoUnderstanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。