arXiv:2503.06934cs.CV2025-03ICCV被引 5

用事件相机提升视觉模型对时间细节的理解能力

LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs

  • 结合帧图像与事件数据,通过交叉注意力融合时空特征
  • 在真实场景数据集上实现毫秒级时间定位精度
  • 适合需要精确时序理解的视频分析任务

大型多模态模型(LMMs)在场景理解方面表现优异,但在细粒度时空推理上因语言与视觉表征对齐不足而受限。现有方法将文本位置和时长映射到基于帧的视频视觉空间,但受时间稀疏性影响,难以实现语言-视觉的时间协调。为此,我们提出LLaFEA(大型语言与帧-事件助手),利用事件相机实现时序密集感知,并引入帧-事件融合机制。该方法采用交叉注意力整合互补的空间与时间特征,再通过自注意力匹配建立全局时空关联。同时,将文本位置和时长标记嵌入融合后的视觉空间,以增强细粒度对齐。统一框架确保了鲁棒的时空坐标对齐,使LMMs可任意位置、任意时间解读场景。此外,我们构建了一个包含真实世界帧-事件数据与坐标指令的数据集,并通过大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suffer from temporal sparsity that limits language-vision temporal coordination. To address this issue, we introduce LLaFEA (Large Language and Frame-Event Assistant) to leverage event cameras for temporally dense perception and frame-event fusion. Our approach employs a cross-attention mechanism to integrate complementary spatial and temporal features, followed by self-attention matching for global spatio-temporal associations. We further embed textual position and duration tokens into the fused visual space to enhance fine-grained alignment. This unified framework ensures robust spatio-temporal coordinate alignment, enabling LMMs to interpret scenes at any position and any time. In addition, we construct a dataset of real-world frames-events with coordinate instructions and conduct extensive experiments to validate the effectiveness of the proposed method.

多模态事件相机时空理解视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。