用思维链增强策略学习,实现私密保护的本地大模型级联决策。
Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning
- 基于思维链引导的策略学习,动态决定是否将任务转交服务器
- 在三个数据集上优于现有方法,兼顾效率与隐私保护
- 适合对数据隐私敏感的移动端大模型应用
大型语言模型(LLMs)因其在真实任务中的优异表现,受到移动端应用的广泛关注。然而,受限于硬件条件,本地部署的LLM性能往往不理想。一种有前景的解决方案是将较弱的本地(设备端)LLM与更强的服务器端LLM进行级联。现有研究主要优化性能与成本的权衡,但实际应用还面临隐私保护等额外需求,尚未得到充分关注。本文提出P³Defer——一种基于思维链(CoT)增强的隐私保护级联决策策略学习框架。该方法突破了传统以置信度或输出概率为基础的级联机制,通过推理过程建模实现更智能的任务转发决策。大量实验在三个基准数据集上验证了P³Defer的有效性与优越性,显著提升级联效率并降低隐私风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have gained significant attention in on-device applications due to their remarkable performance across real-world tasks. However, on-device LLMs often suffer from suboptimal performance due to hardware limitations. A promising solution to this challenge is cascading a weaker local (on-device) LLM with a more powerful server LLM. While existing research on LLM cascade primarily optimizes the performance-cost trade-off, real-world applications impose additional requirements, such as privacy preservation, which remain largely unaddressed. In this work, we move beyond existing confidence- and logit-based LLM cascade methods and propose $\mathbf{P^{3}Defer}$, a novel Chain-of-Thought (CoT)-enhanced \textbf{p}olicy learning framework for \textbf{p}rivacy-\textbf{p}reserved \textbf{defer}ral decision-making. Our approach effectively improves cascade efficiency while mitigating privacy risks. Extensive experiments on three benchmark datasets demonstrate the effectiveness and superiority of $\mathbf{P^{3}Defer}$ over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。