arXiv:2503.07003cs.CL2025-03ICLR被引 10

大模型说一套做一套,真实行为常与表态不一致。

Large Language Models Often Say One Thing and Do Another

  • 构建新评测基准WDCT,对比模型言论与实际行为
  • 多领域测试显示大模型言行严重不一致
  • 仅对齐言论或行为无法改善另一方表现

随着大型语言模型(LLMs)在各类应用中日益重要并面对多元用户群体,确保其可靠且一致的表现变得愈发关键。本文探讨了评估LLM可靠性的一个核心问题:言论与行为之间的一致性。为量化这一一致性,我们开发了一个名为‘言论与行为一致性测试’(Words and Deeds Consistency Test, WDCT)的新评测基准。该基准在多个领域建立了严格对应的言论类与行为类问题,涵盖观点与行动、非伦理价值与行动、伦理价值与行动、理论与应用。评测结果显示,在不同模型和领域中,言论与行为之间普遍存在不一致现象。随后,我们分别进行仅对齐言论或仅对齐行为的实验,观察其对另一方面的效果。实验表明,仅对齐言论或行为对另一方面的影响微弱且不可预测。这支持了我们的假设:驱动大模型选择言论或行为的底层知识并不存在于统一空间中。

原文摘要 · Abstract (English)

As large language models (LLMs) increasingly become central to various applications and interact with diverse user populations, ensuring their reliable and consistent performance is becoming more important. This paper explores a critical issue in assessing the reliability of LLMs: the consistency between their words and deeds. To quantitatively explore this consistency, we developed a novel evaluation benchmark called the Words and Deeds Consistency Test (WDCT). The benchmark establishes a strict correspondence between word-based and deed-based questions across different domains, including opinion vs. action, non-ethical value vs. action, ethical value vs. action, and theory vs. application. The evaluation results reveal a widespread inconsistency between words and deeds across different LLMs and domains. Subsequently, we conducted experiments with either word alignment or deed alignment to observe their impact on the other aspect. The experimental results indicate that alignment only on words or deeds poorly and unpredictably influences the other aspect. This supports our hypothesis that the underlying knowledge guiding LLMs' word or deed choices is not contained within a unified space.

大模型评测言行不一一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。