arXiv:2509.25175cs.CLcs.AI2025-09被引 12

EasySteer让大模型推理时可控更高效,支持多种调控方式。

EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering

  • 模块化设计,支持分析与学习两类调控方法
  • 比现有框架快10.8至22.3倍,支持8类应用的预计算向量
  • 适合需要快速部署可控大模型的研究与工程团队

大语言模型(LLM)控制已成为一种在推理阶段通过操控隐藏状态来调节模型行为的有前景范式,提供了一种轻量级替代昂贵重训练的方法。然而,现有控制框架存在计算效率低、可扩展性差和功能受限等关键问题,阻碍了研究进展与实际部署。我们提出 EasySteer,一个基于 vLLM 的高性能、可扩展统一控制框架。该系统采用模块化架构,支持分析型与学习型方法的即插即用接口,实现细粒度参数控制,并提供针对八类应用领域的预计算控制向量,还配备交互式演示系统。通过与 vLLM 优化推理引擎深度集成,EasySteer 实现了相较于现有框架 10.8–22.3× 的加速。大量实验验证其在过度思考缓解、幻觉减少等关键任务中的有效性。EasySteer 将控制技术从研究工具转变为可部署能力,为可控制、可部署的大语言模型奠定了关键基础设施。

原文摘要 · Abstract (English)

Large language model (LLM) steering has emerged as a promising paradigm for controlling model behavior at inference time through targeted manipulation of hidden states, offering a lightweight alternative to expensive retraining. However, existing steering frameworks suffer from critical limitations: computational inefficiency, limited extensibility, and restricted functionality that hinder both research progress and practical deployment. We present EasySteer, a unified framework for high-performance, extensible LLM steering built on vLLM. Our system features modular architecture with pluggable interfaces for both analysis-based and learning-based methods, fine-grained parameter control, pre-computed steering vectors for eight application domains, and an interactive demonstration system. Through deep integration with vLLM's optimized inference engine, EasySteer achieves 10.8-22.3$\times$ speedup over existing frameworks. Extensive experiments demonstrate its effectiveness in overthinking mitigation, hallucination reduction, and other key applications. EasySteer transforms steering from research technique to production-ready capability, establishing critical infrastructure for deployable, controllable language models.

大模型控制推理优化vLLM可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。