arXiv:2604.11627cs.CV2026-04

让视觉大模型像人一样智能省力,长视频理解快又准。

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

  • 双模式感知:专注模式保精度,待机模式仅用1/40-1/10视觉标记
  • 长视频任务中待机模式保留97.7%-99.7%原始准确率
  • 支持流式输入,动态分离缓存实现超长视觉记忆

多模态大语言模型(MLLMs)在跨模态理解与生成方面表现卓越,但长视频和流式场景下视觉标记序列的快速增长严重制约其可扩展性与实际部署。为此,我们提出POINTS-Long,一种原生双模式MLLM,其视觉标记动态缩放机制受人类视觉系统启发。该模型支持两种互补感知模式:专注模式与待机模式,用户可在推理时动态权衡效率与精度。在细粒度视觉任务中,专注模式保持最优性能;在长时通用视觉理解任务中,待机模式仅使用1/40至1/10的视觉标记,仍能保留97.7%-99.7%的原始准确率。此外,POINTS-Long通过可动态分离的键值缓存设计,原生支持流式视觉理解,实现超长视觉记忆的高效维护。本工作为未来MLLM的设计提供了新视角,奠定了自适应、高效长时视觉理解的基础。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world deployment. Thus, we introduce POINTS-Long, a native dual-mode MLLM featuring dynamic visual token scaling inspired by the human visual system. The model supports two complementary perception modes: focus mode and standby mode, enabling users to dynamically trade off efficiency and accuracy during inference. On fine-grained visual tasks, the focus mode retains the optimal performance, while on long-form general visual understanding, the standby mode retains 97.7-99.7% of the original accuracy using only 1/40-1/10th of the visual tokens. Moreover, POINTS-Long natively supports streaming visual understanding via a dynamically detachable KV-cache design, allowing efficient maintenance of ultra-long visual memory. Our work provides new insights into the design of future MLLMs and lays the foundation for adaptive and efficient long-form visual understanding.

多模态长视频视觉记忆效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。