提出DEMO框架解决视频目标计数中前景背景失衡问题
Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
- 用密度图作为辅助模态实现跨模态自表示学习
- 通过密度引导的自适应掩码提升模型效率与精度
- 构建自然场景下的迁徙鸟类视频数据集DroneBird
视频目标计数中前景与背景动态失衡是主要挑战,通常由目标稀疏性引起,现有研究对此关注不足,常导致严重高估或低估。为此,本文提出一种密度嵌入的高效掩码自编码器计数框架(E-MAC)。为增强密度回归表示能力,设计了新型密度嵌入掩码建模(DEMO)方法,将密度图作为辅助模态,实现图像与密度图的多模态自表示学习。尽管DEMO能有效提供跨模态回归指导,但会引入冗余背景信息,影响对前景的关注。为此,提出基于密度图的高效空间自适应掩码策略以提升效率。同时,采用基于光流的时序协同融合策略,捕捉帧间动态变化,对齐特征并生成多帧密度残差,从而提升当前帧计数精度。此外,针对现有数据集多限于人为主的场景,首次构建了自然场景下用于迁徙鸟类保护的大规模视频鸟类计数数据集DroneBird。在三个行人数据集及DroneBird上的大量实验验证了方法优越性。代码与数据集已开源。
原文摘要 · Abstract (English)
The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model's representation ability on density regression, we develop a new $\mathtt{D}$ensity-$\mathtt{E}$mbedded $\mathtt{M}$asked m$\mathtt{O}$deling ($\mathtt{DEMO}$) method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although $\mathtt{DEMO}$ contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we first propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our \textit{DroneBird} validate our superiority against the counterparts. The code and dataset are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。