提出无人工标注的度量方法,识别大模型偏好学习中的词汇偏差。
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

- 通过三重对比(人类标准、基础模型、指令变体)自动识别偏好学习引入的词汇变化。
- 在六个模型家族中发现偏好学习导致对特定词汇(如delve, furthermore)的过度使用。
- 适用于关注模型对齐与可信AI开发的研究者,尤其适合评估语言风格偏移。
近年来多个语言领域发生显著变化,主要归因于大语言模型与自然语言使用的错位。这种错位部分源于偏好学习阶段(如基于人类反馈的强化学习),该阶段虽提升模型实用性,但可能引入系统性词汇偏差。例如,模型偏好某些表达格式或过度使用特定词汇(如delve, furthermore),即使这些现象未出现在基础模型输出中。现有研究受限于人工标注。本文提出三角化偏好迁移分数(Triangulated Preference Shift score),通过对比人类黄金标准、基础模型和指令变体,无需人工标注即可隔离偏好学习带来的特定行为变化。我们在六个模型家族上提供数据,结合文献锚定结果,并分析偏好学习是否使模型趋向所谓“精英语言”风格。该度量为量化偏好调优引起的可解释行为变化提供了初步自动化方法,有助于模型对齐与可信AI发展。
原文摘要 · Abstract (English)
Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignments are thought to partly originate in the preference-learning stage, e.g. Reinforcement Learning from Human Feedback, which generally makes models more useful but simultaneously may introduce systematic lexical bias. In terms of lexical behavior, this is visible in a model's preference for certain formats or the overuse of words (delve, furthermore), even when such patterns are not present in base model outputs. Research on lexical misalignment induced during preference training is constrained by reliance on manual curation. We address this, by introducing the Triangulated Preference Shift score, a metric that triangulates between human gold standards, base models, and instruct variants to isolate shifts induced specifically by preference learning, without manual curation. We provide data across six model families, anchor the results in the literature, and illustrate the general approach's utility by analyzing whether preference learning shifts models toward what could be interpreted as a "language of prestige". The metric provides an initial automated method to quantify behavioral shifts attributable to preference tuning, and thus, may help inform model alignment and development of trustworthy AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。