arXiv:2505.20088cs.CL2025-05EMNLP被引 2

自动解析大模型偏好背后的多领域概念,让判断逻辑可解释。

Multi-Domain Explainability of Preferences

  • 用大模型识别选择与拒绝响应间的差异概念,生成概念向量。
  • 构建分层多域回归模型,同时捕捉通用与特定领域的偏好影响。
  • 提升大模型输出质量,优化人工与模型判别效果,适合对齐研究者。

偏好机制(如人类偏好、LLM作为裁判、奖励模型)在对齐和评估大语言模型中至关重要,但其背后的驱动概念仍不清晰。本文提出一种全自动方法,实现跨多个领域的局部与全局概念化解释。该方法利用大模型识别区分优选与次选回复的概念,并以概念向量表示;为建模概念与偏好间关系,提出白盒分层多域回归模型,捕捉领域通用与特定效应。我们构建了涵盖八个挑战性多样领域的数据集,解释十二种偏好机制。所提方法在偏好预测上优于基线,且具备可解释性。进一步在两个应用场景中验证:使用LaaJ解释中的概念引导大模型输出,可获得裁判更偏好结果;用解释人类偏好的概念提示LaaJ,能提升其预测性能。本工作确立了大模型时代可解释性的新范式。

原文摘要 · Abstract (English)

Preference mechanisms, such as human preference, LLM-as-a-Judge (LaaJ), and reward models, are central to aligning and evaluating large language models (LLMs). Yet, the underlying concepts that drive these preferences remain poorly understood. In this work, we propose a fully automated method for generating local and global concept-based explanations of preferences across multiple domains. Our method utilizes an LLM to identify concepts that distinguish between chosen and rejected responses, and to represent them with concept-based vectors. To model the relationships between concepts and preferences, we propose a white-box Hierarchical Multi-Domain Regression model that captures both domain-general and domain-specific effects. To evaluate our method, we curate a dataset spanning eight challenging and diverse domains and explain twelve mechanisms. Our method achieves strong preference prediction performance, outperforming baselines while also being explainable. Additionally, we assess explanations in two application-driven settings. First, guiding LLM outputs with concepts from LaaJ explanations yields responses that those judges consistently prefer. Second, prompting LaaJs with concepts explaining humans improves their preference predictions. Together, our work establishes a new paradigm for explainability in the era of LLMs.

偏好解释多领域可解释性大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。