arXiv:2511.05474cs.CV2025-11

用语言提示提升小物体检测精度,兼顾速度与效果。

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

  • 融合BERT与改进的视觉主干网络,实现文本与图像特征对齐
  • COCO2017上达到52.6%平均精度,参数量仅为Transformer模型一半
  • 支持多种高效主干结构,适合资源受限场景

本文提出一种基于语义引导的自然语言与视觉融合方法,用于小物体检测中的跨模态交互。通过将BERT语言模型与基于CNN的并行残差双融合特征金字塔网络(PRB-FPN-Net)结合,引入ELAN、MSP、CSP等创新主干结构,优化特征提取与融合。采用词形还原与微调技术,使文本输入的语义线索与视觉特征精准对齐,显著提升小而复杂目标的检测精度。在COCO和Objects365数据集上的实验表明,该模型在COCO2017验证集上达到52.6%的平均精度(AP),显著优于YOLO-World,且参数量仅为基于Transformer的GLIP模型的一半。多组不同主干结构(ELAN、MSP、CSP)测试进一步证明其在多尺度目标处理上的高效性,具备良好的可扩展性与鲁棒性,适用于资源受限环境。研究展示了自然语言理解与先进主干架构融合的巨大潜力,为小物体检测设定了新的准确率、效率与适应性基准。

原文摘要 · Abstract (English)

This paper introduces a cutting-edge approach to cross-modal interaction for tiny object detection by combining semantic-guided natural language processing with advanced visual recognition backbones. The proposed method integrates the BERT language model with the CNN-based Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN-Net), incorporating innovative backbone architectures such as ELAN, MSP, and CSP to optimize feature extraction and fusion. By employing lemmatization and fine-tuning techniques, the system aligns semantic cues from textual inputs with visual features, enhancing detection precision for small and complex objects. Experimental validation using the COCO and Objects365 datasets demonstrates that the model achieves superior performance. On the COCO2017 validation set, it attains a 52.6% average precision (AP), outperforming YOLO-World significantly while maintaining half the parameter consumption of Transformer-based models like GLIP. Several test on different of backbones such ELAN, MSP, and CSP further enable efficient handling of multi-scale objects, ensuring scalability and robustness in resource-constrained environments. This study underscores the potential of integrating natural language understanding with advanced backbone architectures, setting new benchmarks in object detection accuracy, efficiency, and adaptability to real-world challenges.

小物体检测跨模态融合语义引导高效架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。