用高效强化学习提升大模型中立回答能力,效果优于现有方法。
Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
- 采用参数高效强化学习,仅调整少量参数实现性能跃升。
- 中立性评分从97.06%提升至99.08%,细节支持率增41%。
- 适合关注AI伦理与可控生成的研究者和开发者。
本文表明,参数高效强化学习(PE-RL)是一种高效的训练范式,能显著提升大语言模型(LLMs)在敏感话题上生成中立观点(NPOV)的能力,即提供更丰富、多样且客观的回答。通过对比多种强基线(包括LoRA微调、SFT和RLHF),PE-RL在整体中立性质量上超越最强基线(97.06% → 99.08%),在语言学家识别的关键特征上表现更优:支持性细节出现率从60.25%提升至85.21%,过度简化现象减少率从68.74%降至91.43%。定性分析进一步验证了该结果。此外,评估发现PE-RL的一个关键优势——相比全参数更新方法,其对未见话题具有更强泛化能力。为促进后续研究,本文还发布了数据集SHQ-NPOV,并提供了通过多轮人工同行评审与标注员培训构建此类数据集的方法论。
原文摘要 · Abstract (English)
The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on sensitive topics with a Neutral Point of View (NPOV), i.e. to provide significantly more informative, diverse and impartial answers. This is shown by evaluating PE-RL and multiple strong baselines-including LoRA finetuning (strongest baseline), SFT and RLHF. PE-RL not only improves on overall NPOV quality compared to the strongest baseline ($97.06\%\rightarrow 99.08\%$), but also scores much higher on features linguists identify as key to separating sufficient answers from "great'' answers ($60.25\%\rightarrow 85.21\%$ for presence of supportive details, $68.74\%\rightarrow 91.43\%$ for absence of oversimplification). A qualitative analysis corroborates this. Moreover, our evaluation also finds a key property of PE-RL for this task: unlike methods that update all parameters, it generalises out of topic. Finally, to enable further studies we also release the dataset, SHQ-NPOV, and provide a methodology to create such datasets through iterative rounds of human peer-critique and annotator training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。