让机器追踪目标像人一样智能,能适应复杂变化。
Rethinking Generic Object Tracking Toward Human-Level Perceptual Intelligence

- 用先验知识和几何推理增强目标区分能力
- 在形变、干扰、环境变化下仍保持高追踪精度
- 适合研究视觉感知与智能追踪的学者参考
人类视觉感知的核心在于对外部世界的连续且连贯的理解。通过整合观察与积累的经验,人类视觉系统能够持续适应目标及其环境的变化,同时保持视觉连续性。人类视觉可融合先验知识、空间几何与语义上下文来理解复杂场景及其动态变化。作为计算机视觉的核心问题,视觉目标跟踪旨在使机器感知更接近人类视觉。通用目标跟踪(GOT)要求模型仅凭首帧目标边界框初始化,便需在后续动态视频流中持续定位目标。然而未来事件、观测和现实变化本质上不可预测,因此模型的泛化与在线自适应能力仍是瓶颈。当目标发生严重形变、受复杂干扰、遭遇显著环境变化或属于训练时未见类别时,追踪可靠性会下降。本论文旨在通过一系列方法系统提升追踪模型的目标区分力、鲁棒自适应能力与几何推理能力,缩小机器视觉追踪系统与人类视觉感知之间的差距。
原文摘要 · Abstract (English)
At the heart of human visual perception lies the ability to maintain a continuous and coherent understanding of the external world. By integrating observations with accumulated experience, the human visual system can continuously adapt to variations in both the target and its surrounding environment, while preserving robust visual continuity as scene dynamics evolve. Human vision can therefore integrate prior knowledge, spatial geometry, and semantic context to understand complex scenes and their changes. As a core problem in computer vision, visual object tracking aims to bring machine perception closer to human visual perception. These capabilities are central to the task of Generic Object Tracking (GOT). In this task, a visual tracker is initialized only with the bounding box of an arbitrarily specified target in the first frame, and must continuously localize the target in subsequent dynamic visual streams. However, future events, observations, and real-world variations are inherently unpredictable; therefore, the model's generalization and online adaptation capabilities remain bottlenecks. Tracking reliability can deteriorate when the target undergoes severe deformation, is affected by complex distractors, encounters significant environmental changes, or belongs to a category unseen during training. This dissertation aims to narrow the gap between machine visual tracking systems and human visual perception by proposing a series of methods that systematically enhance the target discrimination, robust adaptation, and geometric reasoning capabilities of tracking models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。