arXiv:2511.21668cs.LGcs.AI2025-11被引 1

提出电信数据采样重要性框架,降低训练开销而不损精度。

Through the telecom lens: Are all training samples important?

  • 通过梯度分析识别样本影响与冗余,动态评估每条数据重要性。
  • 在三个真实电信数据集上,减少数据需求和计算开销,性能保持不变。
  • 适合关注绿色AI、高效训练的电信领域研究者与工程师。

人工智能在电信领域的应用日益广泛,从优化无线接入网到管理用户体验,数据量与训练需求急剧上升。电信数据常具噪声大、高维、存储与标注成本高等特点。尽管AI至关重要,现有流程仍假设所有训练样本同等重要。下一代系统需兼顾准确性、效率与可持续性。本文质疑这一假设,聚焦于分析电信训练中单个样本的作用,评估其对计算与能耗的影响。通过跨训练轮次的样本级梯度分析,识别模型学习中的影响模式与冗余。基于此,提出一种样本重要性框架,可选择性优先处理关键数据,减少计算量而不牺牲精度。在三个真实电信数据集上的实验表明,该方法在保留性能的同时显著降低数据需求与计算开销,推动电信领域可持续AI的发展。

原文摘要 · Abstract (English)

The rise of AI in telecommunications, from optimizing Radio Access Networks to managing user experience, has sharply increased data volumes and training demands. Telecom data is often noisy, high-dimensional, costly to store, process, and label. Despite Ai's critical role, standard workflows still assume all training samples contribute equally. On the other hand, next generation systems require AI models that are accurate, efficient, and sustainable.The paper questions the assumptions of equal importance by focusing on applying and analyzing the roles of individual samples in telecom training and assessing whether the proposed model optimizes computation and energy use. we perform sample-level gradient analysis across epochs to identify patterns of influence and redundancy in model learning. Based on this, we propose a sample importance framework thats electively prioritizes impactful data and reduces computation without compromising accuracy. Experiments on three real-world telecom datasets show that our method [reserves performance while reducing data needs and computational overhead while advancing the goals of sustainable AI in telecommunications.

电信AI样本重要性绿色计算高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。