arXiv:2603.05484cs.CV2026-03被引 8

构建181小时多模态长期理解数据集,解决模型在长时序下的记忆与定位难题。

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

  • 提出递归多模态代理(ReMA),动态管理记忆以更新信念状态。
  • 在月级时间跨度上,性能显著优于现有端到端与代理基线方法。
  • 数据集按天/周/月分层设计,支持长期学习与分布外泛化研究。

尽管视频理解数据集已扩展至小时级时长,但通常由密集拼接的片段构成,与真实日常生活的自然状态不符。为此,我们提出MM-Lifelong数据集,专为多模态长期理解设计,包含181.1小时的视频素材,按天、周、月尺度分层,以捕捉不同时间密度。大量评估揭示当前范式存在两大缺陷:端到端多模态大模型(MLLMs)因上下文饱和出现工作记忆瓶颈,而典型代理基线在稀疏的月级时间线中发生全局定位崩溃。为应对这一挑战,我们提出递归多模态代理(ReMA),通过动态记忆管理迭代更新递归信念状态,在多项任务中表现显著优于现有方法。最后,我们设计了分离时间与领域偏见的数据集划分,为监督学习与分布外泛化研究提供严谨基准。

原文摘要 · Abstract (English)

While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To bridge this gap, we introduce MM-Lifelong, a dataset designed for Multimodal Lifelong Understanding. Comprising 181.1 hours of footage, it is structured across Day, Week, and Month scales to capture varying temporal densities. Extensive evaluations reveal two critical failure modes in current paradigms: end-to-end MLLMs suffer from a Working Memory Bottleneck due to context saturation, while representative agentic baselines experience Global Localization Collapse when navigating sparse, month-long timelines. To address this, we propose the Recursive Multimodal Agent (ReMA), which employs dynamic memory management to iteratively update a recursive belief state, significantly outperforming existing methods. Finally, we establish dataset splits designed to isolate temporal and domain biases, providing a rigorous foundation for future research in supervised learning and out-of-distribution generalization.

多模态长时序代理系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。