arXiv:2412.12596cs.CVstat.AP2024-12被引 12

让多视角模型能识别未知类别,同时解释清楚数据融合过程。

OpenViewer: Openness-Aware Multi-View Learning

  • 用伪未知样本模拟开放环境,提前适应未知数据
  • 通过可解释的网络结构提升多源数据融合透明度
  • 针对已知和未知类别分别优化置信度,增强泛化能力

多视角学习方法通过挖掘多源数据间的关联来提升感知性能,通常依赖预定义类别。但在真实场景部署时面临两大开放性挑战:1)可解释性不足,现有黑箱模型的数据融合机制难以解释;2)泛化能力弱,多数模型无法适应包含未知类别的多视角场景。为此,我们提出OpenViewer——一个具备理论支撑的开放性感知多视角学习框架。该框架首先引入伪未知样本生成机制,高效模拟开放多视角环境并提前适配潜在未知样本;随后设计表达增强型深度展开网络,通过系统构建功能先验映射模块,直观提升可解释性,并提供更透明的多视图数据融合机制;此外,建立感知增强的开集训练策略,通过精准提升已知类别置信度、谨慎抑制未知类别不当置信度,显著增强模型泛化能力。实验表明,OpenViewer在保障已知与未知样本识别性能的同时,有效应对开放性挑战。代码已开源于 https://github.com/dushide/OpenViewer。

原文摘要 · Abstract (English)

Multi-view learning methods leverage multiple data sources to enhance perception by mining correlations across views, typically relying on predefined categories. However, deploying these models in real-world scenarios presents two primary openness challenges. 1) Lack of Interpretability: The integration mechanisms of multi-view data in existing black-box models remain poorly explained; 2) Insufficient Generalization: Most models are not adapted to multi-view scenarios involving unknown categories. To address these challenges, we propose OpenViewer, an openness-aware multi-view learning framework with theoretical support. This framework begins with a Pseudo-Unknown Sample Generation Mechanism to efficiently simulate open multi-view environments and previously adapt to potential unknown samples. Subsequently, we introduce an Expression-Enhanced Deep Unfolding Network to intuitively promote interpretability by systematically constructing functional prior-mapping modules and effectively providing a more transparent integration mechanism for multi-view data. Additionally, we establish a Perception-Augmented Open-Set Training Regime to significantly enhance generalization by precisely boosting confidences for known categories and carefully suppressing inappropriate confidences for unknown ones. Experimental results demonstrate that OpenViewer effectively addresses openness challenges while ensuring recognition performance for both known and unknown samples. The code is released at https://github.com/dushide/OpenViewer.

多视角学习开放集识别可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。