arXiv:2606.01615cs.CVcs.MM2026-06被引 26

用生物反应-扩散模型实现视频与文本的动态对齐,提升关键片段识别能力。

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval

论文配图:Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
图 1 · 摘自论文原文
  • 将视频-文本对齐建模为反应-扩散过程,模拟特征随时间演化
  • 在多个数据集上超越现有方法,显著提升关键片段定位准确率
  • 适合研究跨模态融合、视频理解及生物启发模型的学者

视频-语言模型在关键片段检索和亮点检测等任务中至关重要,但常难以捕捉时序视频与语义文本间的动态非线性交互。现有方法依赖静态交叉注意力或提示调优,无法自适应建模模态间演化的关联,导致对齐效果不佳且泛化能力弱。受系统生物学启发,我们提出反应-扩散多模态融合(RDMF)框架,将视频-语言对齐重新构想为反应-扩散(RD)过程,借鉴阿兰·图灵提出的模式形成原理。在RDMF中,视频特征沿时间扩散以捕获上下文,而文本-视频交互被建模为非线性反应,增强相关特征并抑制噪声,形成类生物系统的涌现模式。基于Gray-Scott RD模型,设计了计算高效的融合模块,并通过图灵不稳定性准则进行稳定性和收敛性分析。该框架理论坚实,结合预训练编码器与DETR风格头部,具备实际可行性。初步实验表明其在识别显著视频片段方面优于现有方法,为视频-语言任务提供新范式。

原文摘要 · Abstract (English)

Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between temporal video sequences and textual semantics. Existing approaches, relying on static cross-attention or prompt-tuning mechanisms, fail to adaptively model the evolving relationships between modalities, leading to suboptimal alignment and limited generalization. Inspired by systems biology, we propose \textbf{Reaction-Diffusion Multimodal Fusion (RDMF)}, a novel framework that reimagines video-language alignment as a reaction-diffusion (RD) process, drawing on the principles of pattern formation introduced by Alan Turing. In RDMF, video features diffuse across time to capture temporal context, while text-video interactions are modeled as non-linear reactions that amplify relevant features and suppress noise, forming emergent patterns akin to biological systems. Leveraging the Gray-Scott RD model, we design a computationally efficient fusion module that integrates video and text representations, supported by rigorous mathematical analysis of stability and convergence using Turing instability criteria. Our framework is theoretically grounded, employing advanced mathematical tools to ensure stable pattern formation, and is practically viable, incorporating standard components like pretrained encoders and DETR-style heads for moment retrieval and saliency prediction. RDMF represents a pioneering interdisciplinary approach, bridging systems biology and multimedia research to address the limitations of conventional multimodal fusion. Preliminary experiments demonstrate its potential to outperform existing methods in identifying salient video moments, offering a new paradigm for video-language tasks.

多模态融合视频理解反应扩散语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。