arXiv:2603.24636cs.LGcs.AI2026-03中稿 · The ACM Web Confer…被引 1

动态融合多空间信息,提升知识图谱事件预测能力

DyMRL: Dynamic Multispace Representation Learning for Multimodal Event Forecasting in Knowledge Graph

  • 引入欧氏、双曲、复数空间构建动态结构特征学习框架
  • 在四个基准数据集上显著超越现有动态与静态方法
  • 适合从事多模态时序建模与知识图谱研究的学者

准确表征多模态知识对现实场景中的事件预测至关重要。然而,现有研究多集中于静态设定,忽视了多模态知识的动态获取与融合。一方面,在知识获取层面,如何捕捉不同模态的时间敏感性,尤其是动态结构模态;现有动态学习方法常局限于异构空间或简单单空间的浅层结构,难以捕捉深层关系感知的几何特征。另一方面,在知识融合层面,如何学习随时间演化的多模态融合特征;基于静态共注意力的方法难以捕捉不同模态对未来的时变历史贡献。为此,我们提出 DyMRL:一种动态多空间表征学习方法,以高效获取和融合多模态时序知识。首先,在知识获取方面,DyMRL 将欧氏、双曲与复数空间中的时序结构特征集成到关系消息传递框架中,学习深度表示,体现人类关联思维、高阶抽象与逻辑推理能力;预训练模型赋予 DyMRL 时间敏感的视觉与语言智能。其次,在知识融合方面,DyMRL 引入先进的双融合-演化注意力机制,对不同时刻的不同模态给予对称且动态的学习权重。为评估 DyMRL 的事件预测性能,我们构建了四个多模态时序知识图谱基准。大量实验表明,DyMRL 在多个指标上优于当前最优的动态单模态与静态多模态基线方法。

原文摘要 · Abstract (English)

Accurate representation of multimodal knowledge is crucial for event forecasting in real-world scenarios. However, existing studies have largely focused on static settings, overlooking the dynamic acquisition and fusion of multimodal knowledge. 1) At the knowledge acquisition level, how to learn time-sensitive information of different modalities, especially the dynamic structural modality. Existing dynamic learning methods are often limited to shallow structures across heterogeneous spaces or simple unispaces, making it difficult to capture deep relation-aware geometric features. 2) At the knowledge fusion level, how to learn evolving multimodal fusion features. Existing knowledge fusion methods based on static coattention struggle to capture the varying historical contributions of different modalities to future events. To this end, we propose DyMRL, a Dynamic Multispace Representation Learning approach to efficiently acquire and fuse multimodal temporal knowledge. 1) For the former issue, DyMRL integrates time-specific structural features from Euclidean, hyperbolic, and complex spaces into a relational message-passing framework to learn deep representations, reflecting human intelligences in associative thinking, high-order abstracting, and logical reasoning. Pretrained models endow DyMRL with time-sensitive visual and linguistic intelligences. 2) For the latter concern, DyMRL incorporates advanced dual fusion-evolution attention mechanisms that assign dynamic learning emphases equally to different modalities at different timestamps in a symmetric manner. To evaluate DyMRL's event forecasting performance through leveraging its learned multimodal temporal knowledge in history, we construct four multimodal temporal knowledge graph benchmarks. Extensive experiments demonstrate that DyMRL outperforms state-of-the-art dynamic unimodal and static multimodal baseline methods.

知识图谱多模态时序预测动态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。