arXiv:2608.15213cs.CVcs.IR2026-08

提出自适应融合与密度路由专家,提升人群计数精度。

DCA-MoE: Spatially Adaptive Cross-Layer Fusion and Density-Routed Experts for Crowd Counting

论文配图:DCA-MoE: Spatially Adaptive Cross-Layer Fusion and Density-Routed Experts for Crowd Counting
图 1 · 摘自论文原文
  • 按位置自适应融合多层特征,动态调整权重。
  • 基于密度分配不同感受野专家,提升局部精度。
  • 适用于复杂场景下人群密度估计,适合计算机视觉研究者。

人群计数需在视角变化、头像大小差异、遮挡和背景杂乱等复杂条件下准确恢复局部密度。尽管现代计数目标提供强空间监督,但许多多级解码器仍采用空间不变的特征融合,并对所有位置使用统一感受野。本文提出DCA-MoE框架,结合空间自适应层融合(SALF)与密度路由多感受野专家(DR-MoE)。SALF在四个对齐的主干特征上预测位置相关的权重;DR-MoE为每个位置分配局部、中程和大范围残差专家的软混合。采用类EBC的头部重建块密度,结合DMCount监督与辅助路由平衡项训练解码器,不更新主干。在NWPU-Crowd验证集上,基于DINOv3 ViT-L/16的最强配置达31.7 MAE和72.2 RMSE;ViT-B/16完整模型为32.2/75.9。跨数据集结果仍不一致,部分基线报告单种子独立选取的最小值。证据支持空间自适应融合与路由的可行性,但更广泛的成对与多种子评估仍有待开展以实现因果归因。

原文摘要 · Abstract (English)

Crowd counting must recover reliable local density under severe variations in perspective, head scale, occlusion, and background clutter. Although modern counting objectives provide strong spatial supervision, many multi-level decoders still use spatially invariant feature fusion and apply one receptive-field pattern to every location. We propose DCA-MoE, a framework that makes both decisions content dependent while retaining a frozen DINOv3 encoder. Spatially Adaptive Layer Fusion (SALF) predicts position-wise weights over four aligned backbone features, and Density-Routed Multi-Receptive-Field Experts (DR-MoE) assigns each location a soft mixture of local, mid-range, and large-context residual experts. An EBC-style head reconstructs block density, while DMCount supervision and an auxiliary routing-balance term train the decoder without updating the backbone. On the NWPU-Crowd validation split, the strongest paired configuration, based on DINOv3 ViT-L/16, obtains 31.7 MAE and 72.2 RMSE; the matched ViT-B/16 full model obtains a paired 32.2/75.9. Cross-dataset results remain mixed, and several component baselines currently report independently selected minima from a single seed. The evidence therefore supports the feasibility of spatially adaptive fusion and routing, while broader paired and multi-seed evaluation remains necessary for causal attribution.

人群计数多专家模型自适应融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。