MCFNet通过协同融合提升细粒度图像分类精度
MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
- 设计跨模态协同融合模块,增强特征表达与语义对齐
- 在多个基准数据集上准确率显著提升,验证方法有效性
- 适合需要高精度分类的多模态视觉任务研究者
多模态信息处理对提升图像分类性能日益重要。然而,不同模态间复杂隐含的依赖关系常使传统方法难以有效捕捉细粒度语义交互,限制其在高精度分类任务中的应用。为此,我们提出一种新型多模态协同融合网络(MCFNet),用于细粒度分类。该架构包含一个正则化集成融合模块,通过模态特定的正则化策略增强模态内特征表示,同时利用混合注意力机制实现精确的语义对齐。此外,引入多模态决策分类模块,通过加权投票框架整合多种损失函数,联合利用模态间相关性与单模态判别特征。在基准数据集上的大量实验与消融研究证明,所提MCFNet框架在分类准确率上持续提升,证实其建模细微跨模态语义的能力。
原文摘要 · Abstract (English)
Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional methods from effectively capturing fine-grained semantic interactions, thereby limiting their applicability in high-precision classification tasks. To address this issue, we propose a novel Multimodal Collaborative Fusion Network (MCFNet) designed for fine-grained classification. The proposed MCFNet architecture incorporates a regularized integrated fusion module that improves intra-modal feature representation through modality-specific regularization strategies, while facilitating precise semantic alignment via a hybrid attention mechanism. Additionally, we introduce a multimodal decision classification module, which jointly exploits inter-modal correlations and unimodal discriminative features by integrating multiple loss functions within a weighted voting paradigm. Extensive experiments and ablation studies on benchmark datasets demonstrate that the proposed MCFNet framework achieves consistent improvements in classification accuracy, confirming its effectiveness in modeling subtle cross-modal semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。