arXiv:2501.06524cs.CV2025-01被引 7

提出新框架解决多视图多标签分类中数据不完整问题

Multi-View Factorizing and Disentangling: A Novel Framework for Incomplete Multi-View Multi-Label Classification

  • 将多视图表征分解为共性与特异性因子,减少冗余
  • 在5个数据集上优于现有方法,提升分类精度
  • 适用于视图或标签缺失的复杂场景,适合实际应用

多视图多标签分类(MvMLC)因广泛的实际应用而受到关注,但视图和标签的不完整性常由数据收集疏漏和人工标注不确定性导致。同时,如何从不同视图中学习既一致又具特异性的鲁棒表示仍是难题。为此,本文提出一种针对不完整多视图多标签分类(iMvMLC)的新框架。该方法将多视图表示分解为独立的视图共性因子与视图特异性因子,并设计图解耦损失以最大限度减少二者间的冗余。进一步地,将共性表示学习分解为三个子目标:(i) 提取跨视图共享信息,(ii) 消除共性表示中的视图内冗余,(iii) 保留任务相关特征。为此,设计了一个鲁棒的任务相关一致性学习模块,结合掩码跨视图预测(MCP)策略与信息论机制,协同学习高质量共性表示。所有模块均能在视图或标签缺失条件下有效运行,具备良好泛化能力。在五个数据集上的大量实验表明,本方法显著优于当前主流方法。

原文摘要 · Abstract (English)

Multi-view multi-label classification (MvMLC) has recently garnered significant research attention due to its wide range of real-world applications. However, incompleteness in views and labels is a common challenge, often resulting from data collection oversights and uncertainties in manual annotation. Furthermore, the task of learning robust multi-view representations that are both view-consistent and view-specific from diverse views still a challenge problem in MvMLC. To address these issues, we propose a novel framework for incomplete multi-view multi-label classification (iMvMLC). Our method factorizes multi-view representations into two independent sets of factors: view-consistent and view-specific, and we correspondingly design a graph disentangling loss to fully reduce redundancy between these representations. Additionally, our framework innovatively decomposes consistent representation learning into three key sub-objectives: (i) how to extract view-shared information across different views, (ii) how to eliminate intra-view redundancy in consistent representations, and (iii) how to preserve task-relevant information. To this end, we design a robust task-relevant consistency learning module that collaboratively learns high-quality consistent representations, leveraging a masked cross-view prediction (MCP) strategy and information theory. Notably, all modules in our framework are developed to function effectively under conditions of incomplete views and labels, making our method adaptable to various multi-view and multi-label datasets. Extensive experiments on five datasets demonstrate that our method outperforms other leading approaches.

多视图学习多标签分类表示学习不完整数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。