不训练即可通过语义特征三阶统计量识别动作相似性
SemanticMoments: Training-Free Motion Similarity via Third Moment Features
- 用预训练语义模型特征的时序高阶统计量捕捉运动动态
- 在合成与真实数据上均超越现有RGB、光流和文本监督方法
- 适合需要快速部署且关注动作语义的视频理解任务
基于语义动作的视频检索是基础但未解决的问题。现有视频表示方法过度依赖静态外观和场景上下文,而非运动动态,这种偏差源于训练数据和目标。传统以运动为中心的输入(如光流)又缺乏高层动作的语义支撑。为揭示这一固有偏差,我们提出SimMotion基准,结合受控合成数据与新的人工标注真实世界数据集。结果显示,现有模型在该基准上表现不佳,常无法区分运动与外观。为此,我们提出SemanticMoments:一种简单、无需训练的方法,通过在预训练语义模型的特征上计算时序统计量(特别是三阶矩),实现运动感知。在多个基准上,SemanticMoments持续优于现有RGB、光流及文本监督方法,表明语义特征空间中的时序统计量可为运动导向的视频理解提供可扩展且符合感知的基底。
原文摘要 · Abstract (English)
Retrieving videos based on semantic motion is a fundamental, yet unsolved, problem. Existing video representation approaches overly rely on static appearance and scene context rather than motion dynamics, a bias inherited from their training data and objectives. Conversely, traditional motion-centric inputs like optical flow lack the semantic grounding needed to understand high-level motion. To demonstrate this inherent bias, we introduce the SimMotion benchmarks, combining controlled synthetic data with a new human-annotated real-world dataset. We show that existing models perform poorly on these benchmarks, often failing to disentangle motion from appearance. To address this gap, we propose SemanticMoments, a simple, training-free method that computes temporal statistics (specifically, higher-order moments) over features from pre-trained semantic models. Across our benchmarks, SemanticMoments consistently outperforms existing RGB, flow, and text-supervised methods. This demonstrates that temporal statistics in a semantic feature space provide a scalable and perceptually grounded foundation for motion-centric video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。