提出统一框架,优化大模型长文本处理的性价比。
The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management

- 将上下文管理建模为兼顾性能、成本与复用的优化问题。
- 实测可降低25%有效令牌消耗,高阶场景下压缩成本超50%。
- 适合关注部署成本的企业、科研与公共部门应用。
大语言模型日益依赖长上下文处理,但扩展上下文窗口带来显著计算与财务开销。现有上下文缩减方法(如检索与记忆压缩)通常独立评估性能与效率,难以系统比较与部署决策。本文提出「效率前沿」框架,将上下文策略选择建模为部署感知的联合优化问题,综合考虑任务性能、令牌成本与预处理复用,通过摊销成本建模实现决策导向分析。在HotpotQA上的实验揭示了不同操作区间的特征与转换边界,表明部署感知优化可在保持性能前提下降低约25%的有效令牌使用量;在高性能场景中,摊销式内存压缩相比全上下文提示可实现超过50%的令牌成本降低。该框架为企业在科研与公共部门中部署可扩展、高效且可持续的大模型系统提供了原则性与实用性基础。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly rely on long-context processing, but expanding context windows introduces substantial computational and financial costs. Existing context reduction approaches, including retrieval and memory compression methods, are typically evaluated using performance and efficiency metrics independently, limiting systematic comparison and deployment-aware decision-making. This paper introduces The Efficiency Frontier, a unified framework for cost--performance optimization in LLM context management. The framework models context strategy selection as a deployment-aware optimization problem that jointly accounts for task performance, token cost, and preprocessing reuse through amortized cost modeling. Unlike existing evaluations that compare methods in isolation, the proposed framework enables decision-oriented analysis. It identifies when different context management strategies become preferable under varying operational conditions. Experiments on HotpotQA reveal distinct operational regimes and transition boundaries between retrieval-based and preprocessing-based strategies. Results show that deployment-aware optimization reduces effective token usage by approximately 25% at comparable performance, enabling more cost-efficient deployment of large language model systems, while amortized memory compression achieves over 50% lower token cost relative to full-context prompting in higher-performance settings. Overall, the proposed framework provides a principled and practical foundation for evaluating and deploying scalable, efficient, and sustainable LLM systems across enterprise, scientific, and public-sector applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。