通过细粒度分析发现大模型生成与人类偏好不一致的深层原因。
Uncovering Factor Level Preferences to Improve Human-Model Alignment
- 提出自动化框架PROFILE,量化人类与模型在因素层面的偏好差异。
- 发现模型生成时与人类偏好偏差大,但在判断任务中表现接近。
- 利用生成-判断差距改进模型对齐,适合希望提升模型可解释性的研究者。
大型语言模型(LLMs)常表现出与人类偏好偏离的倾向,如偏好特定写作风格或产生冗长输出。现有评估方法依赖粗粒度比较且缺乏可解释性,难以识别驱动这些偏差的因素。为此,我们提出PROFILE框架,用于自动发现并测量人类与LLM在因素层面的偏好对齐情况。我们在摘要生成、指令遵循和基于文档的问答三个任务中进行分析,发现显著差异:尽管模型在生成文本时与人类偏好存在较大因素级偏差,但在判别任务中表现出强对齐。我们证明,利用所识别的生成-判别差距可通过多种方式(包括自指导微调)提升模型对齐效果。本工作凸显了因素级分析在揭示隐藏错位中的价值,并提供了实用的框架以改善模型与人类偏好的一致性。
原文摘要 · Abstract (English)
Large language models (LLMs) often exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. While crucial for improvement, identifying the factors driving these misalignments remains challenging due to existing evaluation methods' reliance on coarse-grained comparisons and lack of explainability. To address this, we introduce PROFILE, an automated framework to uncover and measure factor-level preference alignment of humans and LLMs. Using PROFILE, we analyze preference alignment across three key tasks: summarization, instruction-following, and document-based QA. We find a significant discrepancy: while LLMs show poor factor-level alignment with human preferences when generating texts, they demonstrate strong alignment in discrimination tasks. We demonstrate how leveraging the identified generation-discrimination gap can be used to improve LLM alignment through multiple approaches, including fine-tuning with self-guidance. Our work highlights the value of factor-level analysis for identifying hidden misalignments and provides a practical framework for improving LLM-human preference alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。