arXiv:2607.26607cs.SDcs.LG2026-07中稿 · publication in IEE…

通过动态筛选未知类样本,提升少样本音频分类的准确率与抗干扰能力。

Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement

论文配图:Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
图 1 · 摘自论文原文
  • 利用隐空间内点性评分过滤未知类样本,确保原型更新仅依赖已知类数据。
  • 在三个数据集上达到当前最优性能,尤其在未知类比例变化时表现稳定。
  • 适合需要高鲁棒性的少样本音频分类场景,如语音识别中的新类别检测。

少样本开放集音频分类需在仅有少量标注样本的情况下,对已知类进行分类并拒绝未知类样本。传统转导推理虽能利用全部未标注查询样本改进原型估计,但无法区分已知与未知样本,导致原型易受开放集污染。本文提出一种两阶段转导方法,基于冻结的音频编码器:第一阶段为每个查询样本分配隐空间内点性分数,降低疑似未知类样本的影响,使原型更新主要由已知类证据驱动;第二阶段在联合损失函数下优化原型,该损失结合支持集交叉熵、加权条件熵最小化和加权边缘熵最大化,并采用先验自适应自由能得分进行开放集拒绝,其阈值随未知类先验比例动态调整,实现检测与分类解耦。在三个音频数据集上的实验表明,该方法在多种实验条件下均取得当前最优结果。

原文摘要 · Abstract (English)

Few-shot Open-set audio classification requires classifying query samples from known classes with a few labeled support samples while rejecting query samples from unknown classes. Transductive inference jointly observes the full unlabeled query set to improve prototype estimation, yet standard transductive updates do not distinguish known from unknown query samples, leaving prototypes vulnerable to open-set contamination. Drawing on latent-inlierness weighting and decoupled scoring for unknown-class samples, we propose a two-phase transductive method operating over a frozen audio encoder. First, each query sample is assigned a latent inlierness score that down-weights likely unknown-class samples, so that prototype refinement is driven primarily by known-class evidence. The refined prototypes are then directly optimized on a transductive loss combining support cross-entropy, inlierness-weighted conditional entropy minimization, and inlierness-weighted marginal entropy maximization, while open-set rejection uses a prior-adaptive free-energy score that adjusts its threshold with the prior proportion of unknown-class samples, decoupling detection from classification. Experiments on three audio datasets show our method achieves state-of-the-art results for few-shot open-set audio classification under multiple experimental conditions.

少样本学习开放集分类音频识别原型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。