通过生成未见类别数据,提升开放词汇学习中的分布估计精度
Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
- 基于层次语义树和领域信息生成未见类数据
- 在11个数据集上性能比基线最高提升14%
- 适合需要强泛化能力的开放词汇场景
开放词汇学习需建模包含已见类和未见类数据的开放环境分布。现有方法仅用已见类数据估计分布,因缺少未见类信息导致估计误差无法识别。我们理论证明:通过生成未见类数据可有效估计分布,并将误差上界化。基于此,提出新方法:利用层次语义树与从已见类推断的领域信息生成未见类数据,再通过分布对齐算法最大化后验概率,增强泛化能力。11个数据集上的实验表明,该方法相比基线最高提升14%,验证了其有效性与优越性。
原文摘要 · Abstract (English)
Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inherently unidentifiable. Intuitively, learning beyond the seen classes is crucial for distribution estimation to bound the estimation error. We theoretically demonstrate that the distribution can be effectively estimated by generating unseen-class data, through which the estimation error is upper-bounded. Building on this theoretical insight, we propose a novel open-vocabulary learning method, which generates unseen-class data for estimating the distribution in open environments. The method consists of a class-domain-wise data generation pipeline and a distribution alignment algorithm. The data generation pipeline generates unseen-class data under the guidance of a hierarchical semantic tree and domain information inferred from the seen-class data, facilitating accurate distribution estimation. With the generated data, the distribution alignment algorithm estimates and maximizes the posterior probability to enhance generalization in open-vocabulary learning. Extensive experiments on $11$ datasets demonstrate that our method outperforms baseline approaches by up to $14\%$, highlighting its effectiveness and superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。