用可解释信号提升语言模型对人类偏好的理解能力
Preference learning in shades of gray: Interpretable and bias-aware reward modeling for human preferences
- 融合长度、拒绝、毒性等可解释特征增强文本表征
- 最高达0.84的ROC AUC,显著优于基线模型
- 支持安全与语境化表达的公平性分析,适合偏好建模研究者
语言模型中学习人类偏好仍具根本挑战,因奖励建模依赖微妙、主观的对比而非明确标签。本研究评估了十种大语言模型在标准成对偏好设置下的表现,基线性能低于0.74 ROC AUC,凸显任务难度。为此,我们引入可解释信号:响应长度、拒绝指标、毒性分数及提示-响应语义相似度,使模型能显式捕捉有用性、安全性和相关性。所提混合方法在所有模型上均取得一致改进,最高达0.84 ROC AUC,显著提升成对准确率;DeBERTa-v3-Large表现最佳。此外,结合SHAP与LIME实现细粒度可解释性,揭示模型决策依赖情境化安全与支持性表述,而非孤立关键词。进一步分析显示,虽单个特征边际影响弱,但其交互作用显著影响偏好学习。
原文摘要 · Abstract (English)
Learning human preferences in language models remains fundamentally challenging, as reward modeling relies on subtle, subjective comparisons or shades of gray rather than clear-cut labels. This study investigates the limits of current approaches and proposes a feature-augmented framework to better capture the multidimensional nature of human judgment. Using the Anthropic HHRLHF dataset, we evaluate ten diverse large language models LLMs under a standard pairwise preference setting, where baseline performance remains below 0.74 ROC AUC, highlighting the difficulty of the task. To address this, we enrich textual representations with interpretable signals: response length, refusal indicators, toxicity scores and prompt response semantic similarity, enabling models to explicitly capture key aspects of helpfulness, safety and relevance. The proposed hybrid approach yields consistent improvements across all models, achieving up to 0.84 ROC AUC and significantly higher pairwise accuracy, with DeBERTav3Large demonstrating the best performance. Beyond accuracy, we integrate SHAP and LIME to provide fine-grained interpretability, revealing that model decisions depend on contextualized safety and supportive framing rather than isolated keywords. We further analyze bias amplification, showing that while individual features have weak marginal effects, their interactions influence preference learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。