arXiv:2609.08064cs.LG2026-09

让大模型在推理时可调节风险偏好,无需重训。

Risk-Conditioned Fine-Tuning of Large Language Models

论文配图:Risk-Conditioned Fine-Tuning of Large Language Models
图 1 · 摘自论文原文
  • 训练一个能随用户设定调整风险容忍度的统一模型
  • 同一模型在不同风险水平下表现均优于固定风险模型
  • 适合需要灵活控制生成安全性的实际应用

大型语言模型在部署中常面临罕见但严重的有害输出风险。现有风险规避型强化学习人类反馈(Risk-Averse RLHF)通过优化条件风险价值(CVaR)来降低风险,但其模型仅针对固定风险水平训练,无法在推理时动态调整风险厌恶程度。本文提出风险条件化强化学习人类反馈(Risk-Conditioned RLHF),训练单一策略模型,提供连续的风险控制接口,使用户可在不重新训练或部署多个专用模型的前提下,在推理时自由选择不同风险等级。多基准测试结果表明,该单一模型能在推理时适应不同风险水平,实现更灵活、更具备风险意识的模型部署。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.

大模型安全风险控制RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。