用运动模式匹配视频,无需训练数据,跨模态更稳定。
Flow Intelligence: Robust Feature Matching via Temporal Signature Correlation
- 通过像素块的时序运动模式生成特征,不依赖传统关键点。
- 在噪声、错位和跨模态场景下显著优于传统方法。
- 无需预训练,适合实时应用与多领域视频分析。
视频流间的特征匹配是计算机视觉的核心挑战。随着机器人、监控、遥感和医学影像等领域对鲁棒多模态匹配的需求增加,传统基于空间特征的方法在噪声、错位或跨模态数据下失效。现有深度学习方法虽提升了鲁棒性,但仍依赖大量训练数据且计算开销大。本文提出 Flow Intelligence,一种范式革新方法:仅关注时序运动模式,从连续帧中提取像素块的运动签名,并建立视频间的时序相关性。该运动描述符天然具备平移、旋转、尺度不变性,同时在不同成像模态间保持鲁棒性。该方法无需预训练数据,无需空间特征检测,仅凭时序运动即可实现跨模态匹配,在传统方法失效的复杂场景中表现更优。通过利用运动而非外观,实现了多样化环境下的高效实时视频特征匹配。
原文摘要 · Abstract (English)
Feature matching across video streams remains a cornerstone challenge in computer vision. Increasingly, robust multimodal matching has garnered interest in robotics, surveillance, remote sensing, and medical imaging. While traditional rely on detecting and matching spatial features, they break down when faced with noisy, misaligned, or cross-modal data. Recent deep learning methods have improved robustness through learned representations, but remain constrained by their dependence on extensive training data and computational demands. We present Flow Intelligence, a paradigm-shifting approach that moves beyond spatial features by focusing on temporal motion patterns exclusively. Instead of detecting traditional keypoints, our method extracts motion signatures from pixel blocks across consecutive frames and extract temporal motion signatures between videos. These motion-based descriptors achieve natural invariance to translation, rotation, and scale variations while remaining robust across different imaging modalities. This novel approach also requires no pretraining data, eliminates the need for spatial feature detection, enables cross-modal matching using only temporal motion, and it outperforms existing methods in challenging scenarios where traditional approaches fail. By leveraging motion rather than appearance, Flow Intelligence enables robust, real-time video feature matching in diverse environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。