arXiv:2505.18411cs.CLcs.LG2025-05NeurIPS被引 12

构建多模态时间点过程基准,助力理解视频弹幕的时空语义。

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

  • 基于B站弹幕数据构建多模态事件序列,含时间戳、文本与视频帧。
  • 设计多智能体生成问答对,实现跨模态复杂推理任务挑战。
  • 揭示现有模型在多模态时序建模中的显著短板,推动融合进展。

我们提出DanmakuTPPBench,一个面向大语言模型时代的多模态时间点过程(TPP)建模与理解的综合性基准。尽管TPP广泛用于建模时间事件序列,但现有数据集以单模态为主,限制了需联合处理时间、文本与视觉信息的模型发展。为此,DanmakuTPPBench包含两个互补组件:(1) DanmakuTPP-Events,源自Bilibili平台的新型数据集,用户生成的弹幕自然形成带精确时间戳、丰富文本内容和对应视频帧的多模态事件;(2) DanmakuTPP-QA,通过先进大语言模型与多模态大语言模型驱动的多智能体流水线构建,聚焦于复杂的时空-文本-视觉推理问题。我们使用经典TPP模型与最新MLLMs进行广泛评估,揭示当前方法在建模多模态事件动态方面的显著性能差距与局限性。本基准建立强基线,呼吁将TPP建模更深入地融入多模态语言建模体系。项目页面:https://github.com/FRENKIE-CHIANG/DanmakuTPPBench

原文摘要 · Abstract (English)

We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hindering progress in models that require joint reasoning over temporal, textual, and visual information. To address this gap, DanmakuTPPBench comprises two complementary components: (1) DanmakuTPP-Events, a novel dataset derived from the Bilibili video platform, where user-generated bullet comments (Danmaku) naturally form multi-modal events annotated with precise timestamps, rich textual content, and corresponding video frames; (2) DanmakuTPP-QA, a challenging question-answering dataset constructed via a novel multi-agent pipeline powered by state-of-the-art LLMs and multi-modal LLMs (MLLMs), targeting complex temporal-textual-visual reasoning. We conduct extensive evaluations using both classical TPP models and recent MLLMs, revealing significant performance gaps and limitations in current methods' ability to model multi-modal event dynamics. Our benchmark establishes strong baselines and calls for further integration of TPP modeling into the multi-modal language modeling landscape. Project page: https://github.com/FRENKIE-CHIANG/DanmakuTPPBench

多模态时间建模弹幕分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。