用风险规避方法训练大模型,减少有害内容生成。
Risk-Averse Finetuning of Large Language Models
- 引入条件风险价值(CVaR)优化,降低罕见但严重的有害输出
- 在情感修改和毒性缓解任务中显著减少有毒回复
- 适合关注AI安全与内容合规的研究者和开发者
我们针对大语言模型(LLMs)在特定提示下生成负面或有害内容的问题,提出将风险规避原则融入微调过程,以最小化有害输出的发生,尤其是罕见但影响重大的事件。通过优化条件风险价值(CVaR)的风险度量,该方法使模型在保持生成任务有效性的同时,显著降低有害输出概率。在情感修改和毒性缓解任务上的实证评估表明,结合人类反馈的风险规避强化学习(RLHF)能有效促进更安全、更有建设性的在线交流环境。
原文摘要 · Abstract (English)
We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant events. By optimizing the risk measure of Conditional Value at Risk (CVaR), our methodology trains LLMs to exhibit superior performance in avoiding toxic outputs while maintaining effectiveness in generative tasks. Empirical evaluations on sentiment modification and toxicity mitigation tasks demonstrate the efficacy of risk-averse reinforcement learning with human feedback (RLHF) in promoting a safer and more constructive online discourse environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。