让大模型自动切换客观与个性化模式,兼顾准确与贴心。
PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning
- 通过双模式训练,让模型学会客观与个性化两种推理方式。
- 在测试中实现接近无干扰的个性化表现,提升问题解决能力。
- 适合需要兼顾个性与准确性的对话系统应用。
随着用户对大语言模型个性化需求的增加,个性化信息虽能提升交互体验,却可能损害客观性与事实准确性,尤其当其与问题不匹配时。为此,我们提出PersonaDual框架,支持单一模型中通用客观推理与个性化推理的并行,并根据上下文自适应切换模式。该框架首先通过监督微调(SFT)学习两种推理模式,再通过我们提出的DualGRPO强化学习方法优化模式选择。在客观与个性化基准测试中,PersonaDual在保持个性化优势的同时减少干扰,实现近乎无干扰的表现,并更有效地利用有益的个性化信号提升客观问题求解能力。
原文摘要 · Abstract (English)
As users increasingly expect LLMs to align with their preferences, personalized information becomes valuable. However, personalized information can be a double-edged sword: it can improve interaction but may compromise objectivity and factual correctness, especially when it is misaligned with the question. To alleviate this problem, we propose PersonaDual, a framework that supports both general-purpose objective reasoning and personalized reasoning in a single model, and adaptively switches modes based on context. PersonaDual is first trained with SFT to learn two reasoning patterns, and then further optimized via reinforcement learning with our proposed DualGRPO to improve mode selection. Experiments on objective and personalized benchmarks show that PersonaDual preserves the benefits of personalization while reducing interference, achieving near interference-free performance and better leveraging helpful personalized signals to improve objective problem-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。