通过激活函数控制输出熵,显著提升大模型、强化学习与图像分类性能。
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- 用特殊激活函数约束输出采样熵,实现简单高效的熵控制。
- 在大模型、强化学习和图像分类上分别提升37.4%、超30%和0.69%。
- 计算开销低于7%,适合部署于资源受限场景。
我们提出ERA新范式,通过设计特定激活函数,将模型输出的采样熵控制在预设阈值以上。该方法在多个领域展现广泛有效性:1)对大语言模型(LLMs),使Qwen2.5-Math-7B在AIME 2025评测中得分提升37.4%;2)对连续控制强化学习智能体,在HumanoidBench挑战任务上优于SAC等强基线,性能提升超过30%;3)对图像分类,使ResNet-50在ImageNet上的Top-1准确率提升0.69%。所有增益仅带来小于7%的计算开销。本工作验证了输出激活作为熵控制强大工具的潜力,为设计更简单、更鲁棒的算法开辟新方向。
原文摘要 · Abstract (English)
We propose ERA, a new paradigm that constrains the sampling entropy above given thresholds by applying specially designed activations to the outputs of models. Our approach demonstrates broad effectiveness across different domains: 1) for large language models(LLMs), boosting the AIME 2025 score for Qwen2.5-Math-7B by 37.4%; 2) for continuous control reinforcement learning agents, improving performance by more than 30% over strong baselines such as SAC on the challenging HumanoidBench; 3) for image classification, enhancing ImageNet top-1 accuracy by 0.69% for ResNet-50. These gains are achieved with a computational overhead of less than 7%. Our work validates output activation as a powerful tool for entropy control, opening a new direction for designing simpler and more robust algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。