arXiv:2606.03165cs.CLcs.AI2026-06

无需人工标注,自动检测大模型用词偏差及偏好学习影响。

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

论文配图:Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models
图 1 · 摘自论文原文
  • 提出两个无须人工干预的评估指标,识别词汇过度使用与偏好学习影响。
  • 在六类模型上发现'建议''此外''策略'等词被高频使用,且与偏好学习相关。
  • 方法可扩展至非科学语境和多语言场景,适合模型对齐研究者使用。

数字聊天助手(如ChatGPT)的语言可能偏离人类预期(即存在偏差)。现有研究主要聚焦于科学英语,描述了偏差现象及其成因,部分关联到人类偏好学习训练阶段。但现有方法依赖人工标注。本文提出两种无需人工标注、假设极少的评估指标:词汇对齐度量(Lexical Alignment Score),用于识别词汇过度使用;三角化偏好偏移(Triangulated Preference Shift),量化偏好学习导致的偏移程度。基于PubMed摘要生成文本,在六类模型(Falcon、Gemma、Llama、Mistral、OLMo、Yi)上以窗口文档频次进行测量。该方法可自动识别'建议''此外''策略'等过用词汇,并估计其与偏好学习的关联性。结果复现了先前发现,且在不同参数设置、随机种子及新数据上保持稳定。该方法具备良好可扩展性,可用于系统研究科学英语以外的词汇偏差及其跨语言传播,为未来模型对齐改进与成因理解提供支持。

原文摘要 · Abstract (English)

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what divergences occur and, to some extent, why, linking them to the training stage of human preference learning. Yet, existing approaches rely on manual curation. This paper introduces two curation-free, assumption-light evaluation metrics: the Lexical Alignment Score, which identifies lexical overuse, and the Triangulated Preference Shift, which quantifies how much of such shifts can be attributed to human preference learning. Using PubMed abstracts, continuations were generated and measured using windowed document prevalence across six model families (Falcon, Gemma, Llama, Mistral, OLMo, Yi). The procedure identifies, without manual intervention, overused items such as 'suggest', 'additionally', and 'strategy', and estimates their link to preference learning. Our findings replicate prior work and remain stable across parameter settings, random seeds, and evaluation on further data. The approach scales readily and enables systematic study of lexical (mis)alignment beyond Scientific English and across languages, and as such, the metrics have the potential to contribute to improved alignment for future models and understanding of its origins.

模型对齐词汇偏差自动化评估偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。