通过融合语义与视觉信息,提升柔性物体识别准确率。
A Semantic-Enhanced Heterogeneous Graph Learning Method for Flexible Objects Recognition
- 构建异构图模型,关联语义与视觉节点,增强跨模态特征关联。
- 在FDA、FSCW等数据集上实现优于基线的识别性能。
- 适用于形状多变、外观相似的柔性物体识别任务。
柔性物体识别因形状尺寸多样、透明特性及类间差异细微而面临挑战。基于图的模型(如图卷积网络和图视觉模型)因其能捕捉柔性物体内部可变关系而具潜力,但往往只关注全局视觉关系,或未能对齐语义与视觉信息。为此,本文提出一种语义增强的异构图学习方法:首先,采用自适应扫描模块提取判别性语义上下文,实现不同形状尺寸柔性物体的匹配,并对齐语义与视觉节点以增强跨模态特征相关性;其次,设计异构图生成模块,聚合全局视觉与局部语义节点特征,提升识别效果。此外,构建了大规模柔性物体数据集FSCW,源自现有资源。在FDA、FSCW以及挑战性基准CIFAR-100和ImageNet-Hard上的大量实验验证了该方法的竞争力。
原文摘要 · Abstract (English)
Flexible objects recognition remains a significant challenge due to its inherently diverse shapes and sizes, translucent attributes, and subtle inter-class differences. Graph-based models, such as graph convolution networks and graph vision models, are promising in flexible objects recognition due to their ability of capturing variable relations within the flexible objects. These methods, however, often focus on global visual relationships or fail to align semantic and visual information. To alleviate these limitations, we propose a semantic-enhanced heterogeneous graph learning method. First, an adaptive scanning module is employed to extract discriminative semantic context, facilitating the matching of flexible objects with varying shapes and sizes while aligning semantic and visual nodes to enhance cross-modal feature correlation. Second, a heterogeneous graph generation module aggregates global visual and local semantic node features, improving the recognition of flexible objects. Additionally, We introduce the FSCW, a large-scale flexible dataset curated from existing sources. We validate our method through extensive experiments on flexible datasets (FDA and FSCW), and challenge benchmarks (CIFAR-100 and ImageNet-Hard), demonstrating competitive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。