提出新模型与基准,实现文本中时间事件序列的精准检索
Retrieval of Temporal Event Sequences from Textual Descriptions
- 融合大模型与时间点过程,统一编码事件内容与时间信息
- 在多个真实数据集上超越基线模型,显著提升检索准确率
- 适合关注时序分析、行为追踪的应用研究者使用
从文本描述中检索时间事件序列对电商行为分析、社交媒体监控和犯罪事件追踪等应用至关重要。为此,我们提出 TESRBench,一个面向文本描述中时间事件序列检索(TESR)的综合性基准。该基准包含多样化的现实世界数据集,其中的文本描述经过合成与人工审核,为评估检索性能及应对该领域挑战提供了坚实基础。基于此基准,我们提出 TPP-Embedding 模型,通过结合大语言模型(LLMs)与时间点过程(TPPs),同时编码事件文本与发生时间。利用表示池化与对比损失,该模型在共享嵌入空间中统一时间动态与事件语义,实现事件序列与其描述的序列级嵌入对齐。在 TESRBench 各数据集上,TPP-Embedding 均显著优于基线模型,成为该任务的强大解决方案。
原文摘要 · Abstract (English)
Retrieving temporal event sequences from textual descriptions is crucial for applications such as analyzing e-commerce behavior, monitoring social media activities, and tracking criminal incidents. To advance this task, we introduce TESRBench, a comprehensive benchmark for temporal event sequence retrieval (TESR) from textual descriptions. TESRBench includes diverse real-world datasets with synthesized and reviewed textual descriptions, providing a strong foundation for evaluating retrieval performance and addressing challenges in this domain. Building on this benchmark, we propose TPP-Embedding, a novel model for embedding and retrieving event sequences. The model leverages the TPP-LLM framework, integrating large language models (LLMs) with temporal point processes (TPPs) to encode both event texts and times. By pooling representations and applying a contrastive loss, it unifies temporal dynamics and event semantics in a shared embedding space, aligning sequence-level embeddings of event sequences and their descriptions. TPP-Embedding demonstrates superior performance over baseline models across TESRBench datasets, establishing it as a powerful solution for the temporal event sequence retrieval task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。