arXiv:2602.17465cs.CL2026-02

用熵值筛选数据,让语言模型在少数据下高效训练。

Entropy-Based Data Selection for Language Models

  • 基于数据熵值无监督筛选,降低计算开销。
  • 在情感分析等任务上,用更少数据达到相近效果。
  • 适合资源受限场景下的模型快速微调。

现代语言模型日益依赖计算与数据资源。数据筛选技术可减少微调所需数据量,但其效果常受计算资源制约。在实际微调中资源有限的情况下,本文系统揭示了数据筛选与所选数据不确定性估计之间的关系。尽管大语言模型在理解和生成方面表现出色,有助于缓解数据稀缺问题,但评估数据可用性仍具挑战。为此,提出无监督的熵基数据筛选框架(EUDS)。在情感分析(SA)、主题分类(Topic-CLS)和问答(Q&A)任务上的实验证明其有效性。EUDS构建了高效的过滤机制,理论分析与实验结果均证实其可行。该方法显著降低计算成本,提升训练效率,同时减少数据需求,为计算资源受限场景下的语言模型高效微调提供了新方案。

原文摘要 · Abstract (English)

Modern language models (LMs) increasingly require two critical resources: computational resources and data resources. Data selection techniques can effectively reduce the amount of training data required for fine-tuning LMs. However, their effectiveness is closely related to computational resources, which always require a high compute budget. Owing to the resource limitations in practical fine-tuning scenario, we systematically reveal the relationship between data selection and uncertainty estimation of selected data. Although large language models (LLMs) exhibit exceptional capabilities in language understanding and generation, which provide new ways to alleviate data scarcity, evaluating data usability remains a challenging task. This makes efficient data selection indispensable. To mitigate these issues, we propose Entropy-Based Unsupervised Data Selection (EUDS) framework. Empirical experiments on sentiment analysis (SA), topic classification (Topic-CLS), and question answering (Q&A) tasks validate its effectiveness. EUDS establishes a computationally efficient data-filtering mechanism. Theoretical analysis and experimental results confirm the effectiveness of our approach. EUDS significantly reduces computational costs and improves training time efficiency with less data requirement. This provides an innovative solution for the efficient fine-tuning of LMs in the compute-constrained scenarios.

数据筛选语言模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。