用自举方法构建专家角色路由,提升大模型对齐性且不损失准确率。
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
- 通过自蒸馏构建基于意图的专家角色门控适配器。
- 在生成任务中显著提升人类偏好与安全性对齐,判别任务保持高准确率。
- 无需外部数据或模型,内存和计算开销极低,适合部署于多场景。
角色提示可引导大语言模型生成特定领域语调与模式,适用于多智能体系统及需高人类对齐的人类中心任务。然而,现有研究对角色提示效果评价不一:部分显示其在特定领域提升性能并增强合成数据多样性,另一些则发现其对通用能力影响微乎其微甚至为负。本文系统研究了模型优化、任务类型、提示长度与位置对专家角色有效性的影响,揭示其成功与失败的条件。基于此,提出无需外部数据、模型或知识的自举式管道PRISM(基于意图的自我建模角色路由),通过自蒸馏将意图条件化的专家角色嵌入门控LoRA适配器。PRISM在所有模型上均提升了生成任务中的人类偏好与安全对齐,同时保持判别任务的准确性,且仅需极小内存与计算开销。
原文摘要 · Abstract (English)
Persona prompting can steer LLM generation towards a domain-specific tone and pattern. This behavior enables use cases in multi-agent systems where diverse interactions are crucial and human-centered tasks require high-level human alignment. Prior works provide mixed opinions on their utility: some report performance gains when using expert personas for certain domains and their contribution to data diversity in synthetic data creation, while others find near-zero or negative impact on general utility. To fully leverage the benefits of the LLM persona and avoid its harmfulness, a more comprehensive investigation of the mechanism is crucial. In this work, we study how model optimization, task type, prompt length, and placement can impact expert persona effectiveness across instruction-tuned and reasoning LLMs, and provide insight into conditions under which expert personas fail and succeed. Based on our findings, we developed a pipeline to fully leverage the benefits of an expert persona, named PRISM (Persona Routing via Intent-based Self-Modeling), which self-distills an intent-conditioned expert persona into a gated LoRA adapter through a bootstrapping process that requires no external data, models, or knowledge. PRISM enhances human preference and safety alignment on generative tasks while maintaining accuracy on discriminative tasks across all models, with minimal memory and computing overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。