arXiv:2503.04500cs.CVcs.AI2025-03

用流体物理原理分解视频运动,提升动作识别与小物体检测性能

ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding

  • 基于流体动力学理论分解运动为无旋和无散分量
  • 在多个基准上超越或媲美现有方法,计算开销更低
  • 适合需要高效、可解释视频理解的场景

视频理解主要依赖3D卷积网络和光流模型,但这些方法计算成本高,且对光照、尺度和结构变化敏感。为此,我们提出ReynoldsFlow,一种基于雷诺传输定理(RTT)和赫尔姆霍兹-霍奇分解(HHD)的物理启发式时空表示。该方法将运动分解为无旋(CF)和无散(DF)分量,提供可解释的动态表征。结合强度信息与分解后的运动特征,生成具有动态感知且纹理保留的特征,显著提升姿态估计、动作识别和微小目标检测等下游任务表现。ReynoldsFlow轻量且模块化,可无缝集成至现有架构。在多个基准上的实验表明,其性能持续优于或媲美现有方法,兼具更强泛化能力与计算效率。

原文摘要 · Abstract (English)

Video understanding has largely relied on deep spatiotemporal architectures, including 3D convolutional networks and optical flow (OF) based models. While effective, these methods are often computationally expensive and depend on heuristic motion representations that are sensitive to illumination, scale, and structural changes. To address these limitations, we propose ReynoldsFlow, a physics-inspired representation grounded in the Reynolds transport theorem (RTT) and Helmholtz-Hodge decomposition (HHD). ReynoldsFlow decomposes motion into curl-free (CF) and divergence-free (DF) components, providing a principled and interpretable characterization of scene dynamics. By coupling intensity information with decomposed motion cues, it produces dynamics-aware, texture-preserving features that boost downstream tasks such as pose estimation, action recognition, and tiny object detection. Lightweight and modular, ReynoldsFlow can be readily integrated into existing architectures. Experiments across diverse benchmarks show that ReynoldsFlow consistently matches or surpasses existing approaches, offering improved generalizability and computational efficiency.

视频理解流体建模运动分解高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。