arXiv:2512.20021stat.MLcs.LG2025-12

用高斯过程指导数据采集,提升图像分类与目标检测模型性能。

Gaussian Process Assisted Meta-learning for Image Classification and Object Detection Models

  • 基于元学习和高斯过程,根据数据采集条件评估模型表现。
  • 相比随机选择数据,模型在真实场景下的识别准确率提升显著。
  • 适合需要高效采集高质量训练数据的视觉任务研究者。

获取真实操作场景下的数据以训练机器学习模型成本高昂。在收集新数据前,了解模型的薄弱环节至关重要。例如,在稀有物体图像上训练的目标检测器,在低代表性的条件下识别能力较差。本文提出一种方法,通过计算机实验工具和描述训练数据采集条件的元数据(如季节、时间、地点)来指导后续数据采集,以最大化模型性能。具体做法是:按元数据变化调整训练数据,评估模型表现,并用高斯过程拟合其响应曲面,从而指导新的数据采集。该元学习方法在经典学习任务及一个真实案例——搜寻飞机的航拍图像数据采集中均优于随机选择元数据的方案,有效提升了模型性能。

原文摘要 · Abstract (English)

Collecting operationally realistic data to inform machine learning models can be costly. Before collecting new data, it is helpful to understand where a model is deficient. For example, object detectors trained on images of rare objects may not be good at identification in poorly represented conditions. We offer a way of informing subsequent data acquisition to maximize model performance by leveraging the toolkit of computer experiments and metadata describing the circumstances under which the training data was collected (e.g., season, time of day, location). We do this by evaluating the learner as the training data is varied according to its metadata. A Gaussian process (GP) surrogate fit to that response surface can inform new data acquisitions. This meta-learning approach offers improvements to learner performance as compared to data with randomly selected metadata, which we illustrate on both classic learning examples, and on a motivating application involving the collection of aerial images in search of airplanes.

元学习高斯过程数据采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。