arXiv:2604.01015cs.CV2026-04被引 6

用点轨迹建模动物运动,实现复杂行为的精准预测。

Forecasting Motion in the Wild

  • 用密集点轨迹作为运动视觉令牌,分离运动与外观
  • 在300小时野生动物视频上实现跨物种高精度预测
  • 适合需要泛化能力的野外视觉智能研究者

视觉智能需预测代理未来行为,但现有视觉系统缺乏通用运动表征。本文提出以密集点轨迹作为行为的视觉令牌,一种结构化的中层表示,可解耦运动与外观,并在多种非刚性生物(如野外动物)上通用。基于此抽象,设计了扩散变压器,能处理无序轨迹集并显式推理遮挡,实现复杂运动模式的连贯预测。为规模化评估,构建了300小时无约束动物视频数据集,包含鲁棒镜头检测与相机运动补偿。实验表明,轨迹令牌预测实现类别无关、数据高效,优于现有最先进方法,且能泛化至罕见物种与形态,为野外预测视觉智能提供基础。

原文摘要 · Abstract (English)

Visual intelligence requires anticipating the future behavior of agents, yet vision systems lack a general representation for motion and behavior. We propose dense point trajectories as visual tokens for behavior, a structured mid-level representation that disentangles motion from appearance and generalizes across diverse non-rigid agents, such as animals in-the-wild. Building on this abstraction, we design a diffusion transformer that models unordered sets of trajectories and explicitly reasons about occlusion, enabling coherent forecasts of complex motion patterns. To evaluate at scale, we curate 300 hours of unconstrained animal video with robust shot detection and camera-motion compensation. Experiments show that forecasting trajectory tokens achieves category-agnostic, data-efficient prediction, outperforms state-of-the-art baselines, and generalizes to rare species and morphologies, providing a foundation for predictive visual intelligence in the wild.

运动预测扩散模型野外视觉轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。