用旅行商问题重排词元,让加密语言模型生成更稳定
Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- 用旅行商问题优化词元顺序,提升加密状态下文本生成稳定性
- 实验显示该方法有效防止生成坍塌,提升文本连贯性
- 适合关注隐私保护的加密语言模型应用者
随着用户越来越多地使用私密信息与大语言模型交互,安全加密通信变得至关重要。同态加密(HE)通过允许直接在加密数据上进行计算,提供了一种原则性解决方案。尽管已有研究探索了在HE下运行大语言模型的部分方面,但文本生成,尤其是下一个词元预测,仍缺乏足够关注,是实现实际加密交互的关键障碍。本文提出一种基于旅行商问题(TSP)的词元重排策略,并结合后处理步骤以进一步降低近似误差。理论分析和实验结果表明,该方法能有效防止生成坍塌,提升生成文本的连贯性,并在整个过程中保持数据隐私。总体而言,本工作推进了实用且隐私保护的语言模型推理可行性。
原文摘要 · Abstract (English)
As users increasingly interact with large language models (LLMs) using private information, secure and encrypted communication becomes essential. Homomorphic encryption (HE) provides a principled solution by enabling computation directly on encrypted data. Although prior work has explored aspects of running LLMs under HE, the challenge of text generation, particularly next-token prediction, has received limited attention and remains a key obstacle to practical encrypted interaction. In this work, we propose a TSP-based token reordering strategy to address the difficulties of encrypted text generation, together with a post-processing step that further reduces approximation error. Theoretical analysis and experimental results demonstrate that our method prevents collapse, improves coherence in generated text, and preserves data privacy throughout. Overall, our contributions advance the feasibility of practical and privacy-preserving LLM inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。