用相似性分析法量化大模型与人类认知的对齐程度,发现GPT-4o最接近人类。
A Flexible Method for Behaviorally Measuring Alignment Between Human and Artificial Intelligence Using Representational Similarity Analysis
- 通过成对相似性评分,用RSA方法衡量大模型与人类认知的对齐度。
- GPT-4o在文本任务中表现最佳,图像处理能力反而削弱了对齐效果。
- 可识别影响模型行为人性化的提示和超参数,适合评估认知对齐的研究者。
随着大型语言模型(LLMs)承担关键社会决策角色,衡量其与人类认知的对齐程度变得至关重要。为此,我们引入表征相似性分析(Representational Similarity Analysis, RSA),通过成对相似性评分来量化人工智能与人类认知之间的对齐程度。我们在文本与图像模态下测试了多种大语言模型(LLM)和视觉语言模型(VLM)在语义对齐上的表现,发现GPT-4o在群体与个体层面均表现出最强的人类对齐性,尤其在利用其文本处理能力时,无论输入是文本还是图像。然而,所有模型均未能充分捕捉人类个体间的差异,且仅能与个别参与者产生中等程度的对齐。该方法揭示了若干可调控的超参数与提示词,能够引导模型在个体或群体层面表现出更人类化的行为。成对评分与RSA为跨模态(词汇、句子、图像)的高效灵活对齐评估提供了支持,补充了现有基于准确率的基准任务,有助于理解大模型如何编码知识及与人类认知的表征对齐。
原文摘要 · Abstract (English)
As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems represent information and facilitate comparisons with human understanding across diverse tasks. To meet this need, we adapted Representational Similarity Analysis (RSA), a method that uses pairwise similarity ratings to quantify alignment between AIs and humans. We tested this approach on semantic alignment across text and image modalities, measuring how different Large Language and Vision Language Model (LLM and VLM) similarity judgments aligned with human responses at both group and individual levels. GPT-4o showed the strongest alignment with human performance among the models we tested, particularly when leveraging its text processing capabilities rather than image processing, regardless of the input modality. However, no model we studied adequately captured the inter-individual variability observed among human participants, and only moderately aligned with any individual human's responses. This method helped uncover certain hyperparameters and prompts that could steer model behavior to have more or less human-like qualities at an inter-individual or group level. Pairwise ratings and RSA enable the efficient and flexible quantification of human-AI alignment, which complements existing accuracy-based benchmark tasks. We demonstrate the utility of this approach across multiple modalities (words, sentences, images) for understanding how LLMs encode knowledge and for examining representational alignment with human cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。