让大模型对话更像人:用可解释的评分系统精准对齐人类对话特质。
HAL: Inducing Human-likeness in LLMs with Alignment
- 从对比对话数据中提取显式对话特征,合成可解释的评分
- 在聊天竞技场评测中,对齐后模型被认作人类的概率更高
- 适合想提升模型自然度、需可解释对齐的开发者
将语言模型对齐于人类对话特质(如拟人性)仍具挑战性,因其难以定义、测量与优化。现有改进多依赖规模或广泛监督训练,而非针对性对齐。本文提出人类对齐大模型(HAL),通过可解释、数据驱动的奖励信号,实现对对话拟人性的精准对齐。HAL从对比对话数据中提取显式对话特征,融合为紧凑标量得分,并作为透明奖励用于标准偏好优化。该方法适用于不同规模模型,且不影响整体性能。在大规模聊天机器人竞技场式的人类评估中,经HAL对齐的模型更常被视作具有人类特征。由于其基于明确可解释的特质,该框架支持对齐行为的检查与异常诊断。更广泛而言,它展示了以往难以对齐的软性语言特性——如拟人性——如何被量化并以可解释方式实现对齐。
原文摘要 · Abstract (English)
Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised training, rather than targeted alignment. We introduce Human Aligning LLMs (HAL), a framework for aligning language models to conversational human-likeness using an interpretable, data-driven reward. HAL derives explicit conversational traits from contrastive dialogue data, combines them into a compact scalar score, and uses this score as a transparent reward signal for alignment with standard preference optimization methods. Using this approach, we align models of varying sizes without affecting their overall performance. In large-scale Chatbot Arena-style human evaluations, a model aligned with HAL is more frequently perceived as human-like in conversation. Because HAL operates over explicit, interpretable traits, it enables inspection of alignment behavior and diagnosis of unintended effects. More broadly, HAL demonstrates how soft, qualitative properties of language--previously outside the scope for alignment--can be made measurable and aligned in an interpretable and explainable way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。