arXiv:2602.01125cs.CLcs.LG2026-02

让模型同时理解时间、类型和视觉信息,生成长序列的连贯描述。

Long-range Modeling and Processing of Multimodal Event Sequences

  • 通过时间相似性自适应压缩序列,缓解长文本生成的注意力瓶颈。
  • 在DanmakuTPP-QA上超越现有方法,在预测与生成质量上均有提升。
  • 适合需要多模态事件长期推理的场景,如视频分析、社交内容理解。

时间点过程(TPPs)已成为建模异步事件序列的强大工具。尽管近期研究已将TPPs扩展至处理文本信息,但现有方法在生成丰富多模态内容及推理事件动态方面仍受限。主要挑战在于引入多模态数据会显著增加序列长度,阻碍基于注意力的模型生成需长程理解的连贯长文本描述。本文提出一种新框架,将基于大语言模型的TPPs扩展至视觉模态,将文本生成作为核心能力之一,与时间与类型预测并列。通过基于时间相似性的自适应序列压缩机制,减少序列长度的同时保留关键模式。采用两阶段范式:先在压缩序列上预训练,再针对下游任务进行监督微调。大量实验,包括在具有挑战性的DanmakuTPP-QA基准上,表明该方法在预测准确率和生成文本分析质量上均优于现有最先进基线。

原文摘要 · Abstract (English)

Temporal point processes (TPPs) have emerged as powerful tools for modeling asynchronous event sequences. While recent advances have extended TPPs to handle textual information, existing approaches are limited in their ability to generate rich, multimodal content and reason about event dynamics. A key challenge is that incorporating multimodal data dramatically increases sequence length, hindering the ability of attention-based models to generate coherent, long-form textual descriptions that require long-range understanding. In this paper, we propose a novel framework that extends LLM-based TPPs to the visual modality, positioning text generation as a core capability alongside time and type prediction. Our approach addresses the long-context problem through an adaptive sequence compression mechanism based on temporal similarity, which reduces sequence length while preserving essential patterns. We employ a two-stage paradigm of pre-training on compressed sequences followed by supervised fine-tuning for downstream tasks. Extensive experiments, including on the challenging DanmakuTPP-QA benchmark, demonstrate that our method outperforms state-of-the-art baselines in both predictive accuracy and the quality of its generated textual analyses.

多模态时间序列长序列文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。