用标注几何结构增强检测特征,提升模型泛化与可解释性。
Beyond Task-Driven Features for Object Detection
- 通过标注引导的潜在空间生成密集特征图,融合到主干网络
- 在多类数据集上实现更优定位精度与弱监督下的鲁棒性
- 适合关注标注结构、提升模型可解释性的研究者
现代目标检测器学习的任务驱动特征虽优化了最终任务损失,但常捕捉到捷径关联,无法反映底层标注结构,导致在任务定义变化或标注稀疏时,迁移性、可解释性和鲁棒性受限。本文提出一种标注引导的特征增强框架,将嵌入向量注入目标检测主干网络。该方法从标注引导的潜在空间构建密集空间特征图,并与特征金字塔表示融合,影响区域提议和检测头。在野生动物与遥感数据集上的实验评估了多种监督模式下的分类、定位与数据效率。结果表明,该方法能持续提升目标聚焦能力,降低对背景的敏感性,并增强对未见或弱监督任务的泛化性能。研究证明,将特征对齐于标注几何结构,比纯粹任务优化的特征更具意义。
原文摘要 · Abstract (English)
Task-driven features learned by modern object detectors optimize end task loss yet often capture shortcut correlations that fail to reflect underlying annotation structure. Such representations limit transfer, interpretability, and robustness when task definitions change or supervision becomes sparse. This paper introduces an annotation-guided feature augmentation framework that injects embeddings into an object detection backbone. The method constructs dense spatial feature grids from annotation-guided latent spaces and fuses them with feature pyramid representations to influence region proposal and detection heads. Experiments across wildlife and remote sensing datasets evaluate classification, localization, and data efficiency under multiple supervision regimes. Results show consistent improvements in object focus, reduced background sensitivity, and stronger generalization to unseen or weakly supervised tasks. The findings demonstrate that aligning features with annotation geometry yields more meaningful representations than purely task optimized features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。