arXiv:2507.04756cs.CLcs.AI2025-07被引 15

不微调模型,实时个性化生成内容。

CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering

  • 本地小模型计算差异,云端大模型据此调整输出。
  • 仅上传最终生成的词元,保护隐私且节省资源。
  • 适合对响应速度和数据安全要求高的场景。

个性化已成为适应用户在文化、时间与上下文维度上多样化和动态化需求的关键。现有方法通常依赖集中式微调或单一模型内的静态偏好对齐,难以在个人设备的资源与隐私约束下实现既实时又高质量的个性化。为此,我们提出 CoSteer,一种通过解码时协作实现无微调实时个性化的框架。该方法利用上下文感知与上下文无关的本地小型模型之间的逻辑差异,引导云端大型模型进行调整,确保个性化效果的同时保留大模型能力。个性化过程在本地完成,仅需将最终生成的词元发送至云端,兼顾用户上下文保护与系统效率。在多种任务上的广泛实验表明,CoSteer 能生成高质量个性化内容,兼具有效性与计算高效性。结果验证了其在不同模型与环境下的鲁棒性,证实了其在真实场景中的实用价值。

原文摘要 · Abstract (English)

Personalization has become crucial for adapting models to the diverse and evolving needs of users across cultural, temporal, and contextual dimensions. While existing methods often rely on centralized fine-tuning or static preference alignment within a single model, they struggle to achieve both real-time and high-quality personalization under the resource and privacy constraints of personal devices. To address this challenge, we propose CoSteer, a collaborative framework that enables tuning-free, real-time personalization via decoding-time adaptation. By leveraging logit differences between context-aware and context-agnostic local small models, CoSteer steers cloud-based large models, ensuring effective personalization while preserving the large model's capabilities. Personalization is handled locally, with only final tokens sent to the cloud, maintaining both user context and system efficiency. Through extensive experiments across a wide range of tasks, we demonstrate that CoSteer generates high-quality personalized content, ensuring both effectiveness and computational efficiency. Our results highlight its robustness across models and environments, confirming its practical applicability in real-world scenarios.

个性化解码时隐私保护协同推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。