arXiv:2503.04910cs.CLstat.ME2025-03AAAI被引 4

让大模型更懂用户,用真实偏好训练和评估模型。

Maximizing Signal in Human-Model Preference Alignment

  • 分离人类标注中的噪声与有效信号,提升反馈质量。
  • 遵循最佳实践可减少标注分歧,增强判断可信度。
  • 适合关注模型对齐与公平性的研究者和开发者。

强大大语言模型的出现推动了自然语言理解与生成范式的转变。其创造力、流畅表达及高效信息抽象能力虽具价值,但也带来了评估难题。当前快速上市趋势导致团队依赖低成本自动评估,但无法替代人类判断在模型训练与评估中的作用。本文主张:当终端用户需认同模型决策时(如毒性检测或摘要提炼),应基于真实用户偏好数据进行训练与评估。我们通过解析人类反馈在标注与判断任务中的角色,提出分离标注噪声与信号的方法;证明遵循方法论最佳实践可最小化标注分歧,最大化有效信号;并通过两个护栏分类器的人类评估案例,展示如何使模型行为与用户偏好对齐。本文旨在为研究人员和从业者提供整合人类判断的实用指南,尤其在追求准确、无偏且符合用户需求的生成式AI系统时。

原文摘要 · Abstract (English)

The emergence of powerful LLMs has led to a paradigm shift in Natural Language Understanding and Natural Language Generation. The properties that make LLMs so valuable for these tasks -- creativity, ability to produce fluent speech, and ability to quickly and effectively abstract information from large corpora -- also present new challenges to evaluating their outputs. The rush to market has led teams to fall back on quick, cost-effective automatic evaluations which offer value, but do not obviate the need for human judgments in model training and evaluation. This paper argues that in cases in which end users need to agree with the decisions made by ML models -- e.g. in toxicity detection or extraction of main points for summarization -- models should be trained and evaluated on data that represent the preferences of those users. We support this argument by explicating the role of human feedback in labeling and judgment tasks for model training and evaluation. First, we propose methods for disentangling noise from signal in labeling tasks. Then we show that noise in labeling disagreement can be minimized by adhering to proven methodological best practices, while signal can be maximized to play an integral role in model training and evaluation tasks. Finally, we illustrate best practices by providing a case study in which two guardrails classifiers are evaluated using human judgments to align final model behavior to user preferences. We aim for this paper to provide researchers and professionals with guidelines to integrating human judgments into their ML and generative AI evaluation toolkit, particularly when working toward achieving accurate and unbiased features that align with users' needs and expectations.

人机对齐模型评估人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。