融合有标签源数据与弱监督袋标签,提升目标域预测性能
Learning from Label Proportions and Covariate-shifted Instances
- 在领域自适应框架下联合利用源域实例标签和目标域袋标签
- 在多个公开数据集上优于传统LLP与领域自适应基线方法
- 适合存在标注数据偏移但需跨域泛化的实际场景
在许多应用场景中,由于缺乏监督或隐私限制,训练数据以实例袋的形式组织,每袋仅提供基于袋内实例标签的聚合标签。在学习标签比例(LLP)任务中,聚合标签为袋内实例标签的均值。然而实践中,可能存在已标注但分布偏移的源数据,以及仅有袋标签的目标数据。本文提出一种混合式协变量偏移下的LLP问题求解方法,将源域实例标签与目标域袋标签统一纳入领域自适应框架进行建模。理论分析给出了目标域泛化误差的界,实验在多个公开数据集上验证了该方法在预测性能上优于现有基线及相关工作。
原文摘要 · Abstract (English)
In many applications, especially due to lack of supervision or privacy concerns, the training data is grouped into bags of instances (feature-vectors) and for each bag we have only an aggregate label derived from the instance-labels in the bag. In learning from label proportions (LLP) the aggregate label is the average of the instance-labels in a bag, and a significant body of work has focused on training models in the LLP setting to predict instance-labels. In practice however, the training data may have fully supervised albeit covariate-shifted source data, along with the usual target data with bag-labels, and we wish to train a good instance-level predictor on the target domain. We call this the covariate-shifted hybrid LLP problem. Fully supervised covariate shifted data often has useful training signals and the goal is to leverage them for better predictive performance in the hybrid LLP setting. To achieve this, we develop methods for hybrid LLP which naturally incorporate the target bag-labels along with the source instance-labels, in the domain adaptation framework. Apart from proving theoretical guarantees bounding the target generalization error, we also conduct experiments on several publicly available datasets showing that our methods outperform LLP and domain adaptation baselines as well techniques from previous related work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。