arXiv:2510.26202cs.CLcs.AI2025-10被引 26

用稀疏自编码器解析人类反馈数据,揭示真实偏好与潜在风险。

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

  • 通过稀疏自编码器自动挖掘反馈中的可解释特征
  • 7个数据集均发现少数关键特征主导偏好预测信号
  • 可识别安全风险并支持个性化定制与数据优化

人类反馈常导致语言模型产生不可预测的不良行为,因从业者难以理解反馈数据的真实含义。现有研究多聚焦特定属性(如文本长度或奉承倾向),但无需预设假设即可自动提取相关特征仍具挑战。本文提出WIMHF方法,利用稀疏自编码器解析反馈数据,揭示数据集所能测量的偏好类型及标注者实际表达的偏好。在7个数据集中,该方法识别出少量人类可理解的特征,覆盖了黑箱模型偏好预测信号的大部分。这些特征显示人类偏好存在广泛差异:例如Reddit用户偏爱非正式和幽默,而HH-RLHF与PRISM标注者则相反。此外,模型还暴露潜在危险偏好——如LMArena用户倾向于反对拒绝回答,常支持有害内容。基于学习到的特征,重新标注有害样本可实现+37%的安全提升,且不影响整体性能;同时在社区对齐数据集上,可学习标注者个性化的主观特征权重,显著提升偏好预测效果。WIMHF为从业者提供了以人为本的偏好数据解析工具。

原文摘要 · Abstract (English)

Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies preferences over certain attributes (e.g., length or sycophancy), automatically extracting relevant features without pre-specifying hypotheses remains challenging. We introduce What's In My Human Feedback? (WIMHF), a method to explain feedback data using sparse autoencoders. WIMHF characterizes both (1) the preferences a dataset is capable of measuring and (2) the preferences that the annotators actually express. Across 7 datasets, WIMHF identifies a small number of human-interpretable features that account for the majority of the preference prediction signal achieved by black-box models. These features reveal a wide diversity in what humans prefer, and the role of dataset-level context: for example, users on Reddit prefer informality and jokes, while annotators in HH-RLHF and PRISM disprefer them. WIMHF also surfaces potentially unsafe preferences, such as that LMArena users tend to vote against refusals, often in favor of toxic content. The learned features enable effective data curation: re-labeling the harmful examples in Arena yields large safety gains (+37%) with no cost to general performance. They also allow fine-grained personalization: on the Community Alignment dataset, we learn annotator-specific weights over subjective features that improve preference prediction. WIMHF provides a human-centered analysis method for practitioners to better understand and use preference data.

偏好学习可解释性数据清洗安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。