构建首个面向复杂运动的视频伪装目标检测基准,提出稳定特征与轨迹对齐新方法。
YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection
- 设计动态感知的特征稳定模块和轨迹引导对齐机制
- 在91个场景、47类物种上实现显著性能提升,跨域泛化能力更强
- 适合关注复杂运动下目标检测的科研与工程人员
视频伪装目标检测(VCOD)受限于挑战性数据集稀缺和模型对复杂运动的鲁棒性不足。现有方法常受运动引起的外观不稳定和时序特征错位影响。为此,我们提出YUV20K,一个像素级标注的复杂度驱动型VCOD基准,包含91个场景、47种生物的24,295帧标注数据,专门覆盖大位移运动、相机运动等5类复杂场景。方法上,提出融合运动特征稳定(MFS)与轨迹感知对齐(TAA)的新框架:MFS利用无帧依赖的语义基元稳定特征,TAA通过轨迹引导可变形采样实现精准时序对齐。大量实验表明,该方法在现有数据集上超越当前最优模型,并在挑战性的YUV20K上建立新基线,展现出优异的跨域泛化与复杂时空场景鲁棒性。代码与数据集将开源。
原文摘要 · Abstract (English)
Video Camouflaged Object Detection (VCOD) is currently constrained by the scarcity of challenging benchmarks and the limited robustness of models against erratic motion dynamics. Existing methods often struggle with Motion-Induced Appearance Instability and Temporal Feature Misalignment caused by complex motion scenarios. To address the data bottleneck, we present YUV20K, a pixel-level annoated complexity-driven VCOD benchmark. Comprising 24,295 annotated frames across 91 scenes and 47 kinds of species, it specifically targets challenging scenarios like large-displacement motion, camera motion and other 4 types scenarios. On the methodological front, we propose a novel framework featuring two key modules: Motion Feature Stabilization (MFS) and Trajectory-Aware Alignment (TAA). The MFS module utilizes frame-agnostic Semantic Basis Primitives to stablize features, while the TAA module leverages trajectory-guided deformable sampling to ensure precise temporal alignment. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art competitors on existing datasets and establishes a new baseline on the challenging YUV20K. Notably, our framework exhibits superior cross-domain generalization and robustness when confronting complex spatiotemporal scenarios. Our code and dataset will be available at https://github.com/K1NSA/YUV20K
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。