通过跨尺度上下文融合,提升多标签图像识别准确率。
Multi-label Classification with Panoptic Context Aggregation Networks
- 在高维希尔伯特空间中分层聚合多阶几何关系
- 在NUS-WIDE等数据集上超越现有方法,最高提升3.2%精度
- 适合复杂场景下需要细粒度语义理解的任务
上下文建模对视觉识别至关重要,能通过整合图像中对象与标签之间的内在和外在关系,生成具有强区分性的图像表示。当前方法多聚焦于基础几何关系或局部特征,常忽视对象间的跨尺度上下文交互。本文提出深度全景上下文聚合网络(PanCAN),通过在高维希尔伯特空间中进行跨尺度特征聚合,层次化集成多阶几何上下文。具体而言,PanCAN在每一尺度上结合随机游走与注意力机制学习多阶邻域关系;不同尺度模块级联,细粒度显著锚点被选中,其邻域特征通过注意力动态融合,实现有效的跨尺度建模,显著增强复杂场景理解能力。在NUS-WIDE、PASCAL VOC2007和MS-COCO基准上的大量多标签分类实验表明,PanCAN持续取得竞争力结果,在定量与定性评估中均优于先进方法,显著提升多标签分类性能。
原文摘要 · Abstract (English)
Context modeling is crucial for visual recognition, enabling highly discriminative image representations by integrating both intrinsic and extrinsic relationships between objects and labels in images. A limitation in current approaches is their focus on basic geometric relationships or localized features, often neglecting cross-scale contextual interactions between objects. This paper introduces the Deep Panoptic Context Aggregation Network (PanCAN), a novel approach that hierarchically integrates multi-order geometric contexts through cross-scale feature aggregation in a high-dimensional Hilbert space. Specifically, PanCAN learns multi-order neighborhood relationships at each scale by combining random walks with an attention mechanism. Modules from different scales are cascaded, where salient anchors at a finer scale are selected and their neighborhood features are dynamically fused via attention. This enables effective cross-scale modeling that significantly enhances complex scene understanding by combining multi-order and cross-scale context-aware features. Extensive multi-label classification experiments on NUS-WIDE, PASCAL VOC2007, and MS-COCO benchmarks demonstrate that PanCAN consistently achieves competitive results, outperforming state-of-the-art techniques in both quantitative and qualitative evaluations, thereby substantially improving multi-label classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。