arXiv:2511.20662cs.CL2025-11AAAI

让普通机构也能轻松用上高效大模型。

Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability

  • 不重训模型,直接改造架构提升效率
  • 轻量微调保持模型对齐,推理更经济
  • 关注部署成本与可持续性,适合中小机构

大语言模型虽已不可或缺,但主流高效方法如混合专家(MoE)、推测解码和复杂检索增强生成(RAG)仅适用于拥有海量资源和顶尖团队的超大规模企业。在其他场景下,这些方法反而带来额外开销、脆弱性和碳排放浪费。结果是少数科技巨头获益,而成千上万的医院、学校、政府和企业缺乏可行方案。我们主张,下一个前沿不是更大规模的复杂优化,而是鲁棒的简单性:在有限资源和低门槛下依然高效的模型。提出新研究方向:无需重训即可改造预训练模型架构;开发保持对齐的轻量微调;使长链推理更经济;实现无需重型RAG的动态知识管理;以“开销感知效率”(OAE)为标准评估。通过将效率纳入采纳成本、可持续性与公平性,真正实现大模型部署的民主化,让优化减少不平等与碳足迹,而非加剧。

原文摘要 · Abstract (English)

Large language models (LLMs) have become indispensable, but the most celebrated efficiency methods -- mixture-of-experts (MoE), speculative decoding, and complex retrieval-augmented generation (RAG) -- were built for hyperscale providers with vast infrastructure and elite teams. Outside that context, their benefits collapse into overhead, fragility, and wasted carbon. The result is that a handful of Big Tech companies benefit, while thousands of hospitals, schools, governments, and enterprises are left without viable options. We argue that the next frontier is not greater sophistication at scale, but robust simplicity: efficiency that thrives under modest resources and minimal expertise. We propose a new research agenda: retrofitting pretrained models with more efficient architectures without retraining, inventing lightweight fine-tuning that preserves alignment, making reasoning economical despite long chains of thought, enabling dynamic knowledge management without heavy RAG pipelines, and adopting Overhead-Aware Efficiency (OAE) as a standard benchmark. By redefining efficiency to include adoption cost, sustainability, and fairness, we can democratize LLM deployment -- ensuring that optimization reduces inequality and carbon waste rather than amplifying them.

模型效率普惠AI轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。