用人类相似性判断评估大模型指令引导效果,发现提示法更贴近人脑认知。
Evaluating Steering Techniques using Human Similarity Judgments
- 基于人类认知的三元相似性任务评估模型对概念的灵活判断能力
- 提示法在引导准确性和人机一致性上均优于其他方法
- 模型天然偏好'种类'相似性,难掌握'大小'相似性,揭示潜在表征轴
当前大语言模型(LLM)的指令引导评估多聚焦于特定任务表现,忽视了引导后表征与人类认知的一致性。我们采用经典的三元相似性判断任务,评估了不同引导方式下模型对概念间基于‘大小’或‘种类’的相似性判断能力。结果表明,提示法在引导准确性和模型与人类判断的一致性上均显著优于其他方法。此外,模型普遍倾向于‘种类’相似性,难以实现‘大小’相似性的对齐。该基于人类认知的评估框架进一步验证了提示法的有效性,并揭示了模型在引导前就存在的特权表征轴。
原文摘要 · Abstract (English)
Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. Using a well-established triadic similarity judgment task, we assessed steered LLMs on their ability to flexibly judge similarity between concepts based on size or kind. We found that prompt-based steering methods outperformed other methods both in terms of steering accuracy and model-to-human alignment. We also found LLMs were biased towards 'kind' similarity and struggled with 'size' alignment. This evaluation approach, grounded in human cognition, adds further support to the efficacy of prompt-based steering and reveals privileged representational axes in LLMs prior to steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。