测试大模型对形容词名词组合的理解能力,发现表现与内部状态不一致。
Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
- 用提示词功能测试和内部表征分析双角度评估
- 模型内部有组合性表征但任务表现不稳定
- 适合关注模型真实理解力的研究者阅读
组合性被视为语言能力的核心。作为高性能语言系统,大语言模型(LLMs)在组合性任务上的表现如何?我们通过两种互补的实验设置评估了LLMs在形容词-名词组合上的表现:基于提示的功能性评估和对内部模型状态的表征分析。结果揭示了任务表现与内部状态之间显著的分歧:尽管LLMs能够可靠地形成组合性表征,但在不同模型变体间,这些表征未能一致转化为功能性任务的成功。因此,我们强调对比评估的重要性,以获得对模型能力更完整的理解。
原文摘要 · Abstract (English)
Compositionality is considered central to language abilities. As performant language systems, how do large language models (LLMs) do on compositional tasks? We evaluate adjective-noun compositionality in LLMs using two complementary setups: prompt-based functional assessment and a representational analysis of internal model states. Our results reveal a striking divergence between task performance and internal states. While LLMs reliably develop compositional representations, they fail to translate consistently into functional task success across model variants. Consequently, we highlight the importance of contrastive evaluation for obtaining a more complete understanding of model capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。