arXiv:2603.08145cs.LGcs.AI2026-03

让AI生成更少分歧,用风险控制选更好回答。

DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding

  • 不重新训练,在推理时通过风险约束重排序回复。
  • 在嘈杂反馈下降低分歧与极端风险,平均质量仍领先。
  • 适合需要稳定输出的高敏感场景,如医疗、金融对话。

基于偏好对齐的方法(如RLHF、DPO)通常优化单一标量目标,隐式地对异质人类偏好取均值。现实中,标注者和用户群体间的系统性分歧使得均值奖励最大化方法脆弱且易受代理过优化影响。我们提出**基于风险约束解码的分歧感知对齐方法(DARC)**,一种无需重训练的推理时方法,将响应选择建模为分布鲁棒、风险敏感的决策过程。给定多个偏好样本或可扩展的分歧代理,DARC通过最大化*KL-鲁棒(熵型)满意度目标*对候选回复重排序,并提供简单的部署控制机制,限制或惩罚相对于均值的熵风险溢价,实现无需重训练的显式风险预算。我们提供了理论分析,揭示该解码规则与原则性悲观主义及基于KL的分布鲁棒优化之间的联系。在对齐基准上的实验表明,DARC在噪声和异质反馈下显著降低了分歧与尾部风险,同时保持了有竞争力的平均质量。

原文摘要 · Abstract (English)

Preference-based alignment methods (e.g., RLHF, DPO) typically optimize a single scalar objective, implicitly averaging over heterogeneous human preferences. In practice, systematic annotator and user-group disagreement makes mean-reward maximization brittle and susceptible to proxy over-optimization. We propose **Disagreement-Aware Alignment via Risk-Constrained Decoding (DARC)**, a retraining-free inference-time method that frames response selection as distributionally robust, risk-sensitive decision making. Given multiple preference samples or scalable disagreement proxies, DARC reranks candidates by maximizing a *KL-robust (entropic)* satisfaction objective, and provides simple deployment controls that cap or penalize the corresponding entropic risk premium relative to the mean, enabling explicit risk budgets without retraining. We provide theoretical characterization linking this decoding rule to principled pessimism and KL-based distributionally robust optimization. Experiments on alignment benchmarks show that DARC reduces disagreement and tail risk while maintaining competitive average quality under noisy, heterogeneous feedback.

对齐方法风险控制推理优化多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。