arXiv:2501.02004cs.LGcs.AI2025-01被引 9

用信息论指标选数据,降本增效不降性能。

General Information Metrics for Improving AI Model Training Efficiency

  • 基于客观信息论的11个指标筛选训练数据
  • 在多个领域实现训练成本降低超39%
  • 适合需要高效训练的落地型AI项目

针对人工智能模型训练数据规模增长与缺乏通用数据选择方法导致训练成本攀升的问题,本文提出通用信息度量评估(GIME)方法。GIME基于客观信息论(OIT)中的体积、延迟、范围、粒度、多样性、持续时间、采样率、聚合、覆盖、失真和偏差等11项通用信息度量,优化训练数据集选择。在点击率预测、民事案件预测和天气预报等多个领域开展的全面实验表明,GIME在显著降低训练时间和成本的同时,有效保持了模型性能。将GIME应用于司法AI项目后,总训练开支减少39.56%,展现了其在推动高效可持续人工智能发展方面的潜力。

原文摘要 · Abstract (English)

To address the growing size of AI model training data and the lack of a universal data selection methodology-factors that significantly drive up training costs -- this paper presents the General Information Metrics Evaluation (GIME) method. GIME leverages general information metrics from Objective Information Theory (OIT), including volume, delay, scope, granularity, variety, duration, sampling rate, aggregation, coverage, distortion, and mismatch to optimize dataset selection for training purposes. Comprehensive experiments conducted across diverse domains, such as CTR Prediction, Civil Case Prediction, and Weather Forecasting, demonstrate that GIME effectively preserves model performance while substantially reducing both training time and costs. Additionally, applying GIME within the Judicial AI Program led to a remarkable 39.56% reduction in total model training expenses, underscoring its potential to support efficient and sustainable AI development.

数据筛选训练效率信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。