arXiv:2601.20255cs.LGcs.CL2026-01中稿 · ICML被引 2

用熵压缩理论提升模型训练效果,精准指导复杂编程任务的模型优化

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

  • 提出基于熵压缩的新指标HE-SNR,捕捉模型对不确定性的结构化处理能力
  • 在32K/128K上下文窗口下,560B参数模型的推理性能显著提升
  • 适合关注大模型中段训练与复杂工程任务优化的研究者和工程师

SWE-bench已成为评估大语言模型在复杂软件工程任务中表现的首选基准。尽管这些能力主要在中段训练阶段形成,并在监督微调(SFT)阶段被激活,但目前仍缺乏有效指导中段训练的指标。标准指标如困惑度(PPL)受“长上下文税”影响,与下游SWE表现相关性弱。本文首先提出严格的过滤策略,关键在于提出熵压缩假说:智能不在于单一的Top-1压缩,而在于将不确定性结构化为低阶熵压缩状态(“合理犹豫”)。基于此细粒度熵分析,我们构建新指标HE-SNR(高熵信噪比)。在支持32K/128K上下文、最大达560B参数的模型上验证该方法。本工作为优化大模型在复杂工程领域中的潜在能力提供了理论基础与实用工具。

原文摘要 · Abstract (English)

SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabilities are fundamentally acquired during the mid-training phase and subsequently elicited during Supervised Fine-Tuning (SFT), there remains a critical deficit in metrics capable of guiding mid-training effectively. Standard metrics such as Perplexity (PPL) are compromised by the "Long-Context Tax" and exhibit weak correlation with downstream SWE performance. In this paper, we bridge this gap by first introducing a rigorous data filtering strategy. Crucially, we propose the Entropy Compression Hypothesis, redefining intelligence not by scalar Top-1 compression, but by the capacity to structure uncertainty into Entropy-Compressed States of low orders ("reasonable hesitation"). Grounded in this fine-grained entropy analysis, we formulate a novel metric, HE-SNR (High-Entropy Signal-to-Noise Ratio). We validate our approach on models with up to 560B parameters across different context windows (32K/128K). This work provides both the theoretical foundation and practical tools for optimizing the latent potential of LLMs in complex engineering domains.

大模型训练熵分析编程推理指标设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。