提升多模态大模型对周期性现象的识别与推理能力
Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
- 采用从易到难渐进式训练,逐步构建周期推理能力
- 提出抗逻辑遗忘策略,稳定跨模态周期理解性能
- 构建多难度基准测试,涵盖文本、视觉等多模态周期任务
周期或准周期现象在气象、交通流、生物信号等多种自然过程中普遍存在。由于这些现象涉及多模态数据,多模态大语言模型(MLLM)具备捕捉其复杂特性的潜力。然而,现有MLLM在周期性任务上表现不佳,主要受限于:1)缺乏时间建模能力;2)短周期与长周期之间的冲突。本文提出Period-LLM,一种增强周期性任务处理能力的多模态大模型,并构建了一个涵盖多种难度的跨模态周期性评估基准。特别地,采用“由易到难泛化”范式,从简单的文本任务逐步过渡到复杂的视觉与多模态任务,确保模型逐步建立稳健的周期推理能力。此外,提出“抵抗逻辑遗忘”优化策略,以在语义对齐过程中保持周期推理能力。大量实验表明,Period-LLM在周期性任务上显著优于现有MLLM。代码已公开于https://github.com/keke-nice/Period-LLM。
原文摘要 · Abstract (English)
Periodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) offer promising potential to effectively capture and understand their complex nature. However, current MLLMs struggle with periodic tasks due to limitations in: 1) lack of temporal modelling and 2) conflict between short and long periods. This paper introduces Period-LLM, a multimodal large language model designed to enhance the performance of periodic tasks across various modalities, and constructs a benchmark of various difficulty for evaluating the cross-modal periodic capabilities of large models. Specially, We adopt an "Easy to Hard Generalization" paradigm, starting with relatively simple text-based tasks and progressing to more complex visual and multimodal tasks, ensuring that the model gradually builds robust periodic reasoning capabilities. Additionally, we propose a "Resisting Logical Oblivion" optimization strategy to maintain periodic reasoning abilities during semantic alignment. Extensive experiments demonstrate the superiority of the proposed Period-LLM over existing MLLMs in periodic tasks. The code is available at https://github.com/keke-nice/Period-LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。