arXiv:2504.02304cs.CL2025-04

用心理量表测大模型对人性的看法,发现越聪明越不信任人

Measurement of LLM's Philosophies of Human Nature

  • 基于心理学量表设计新工具,评估大模型对人性的六维度态度
  • 实测显示当前大模型普遍不信任人类,且智能越高越不信任
  • 提出心智循环学习框架,可让模型在虚拟互动中改善对人的态度

人工智能在各领域广泛应用,伴随其与人类冲突频发,引发社会对其交互行为的担忧。基于长期验证的心理学量表Wrightsman人性观量表(PHNS),我们构建了针对大语言模型(LLM)的标准化心理量表——机器人性观量表(M-PHNS)。通过六个维度评估,发现当前大模型普遍存在对人类的系统性不信任,且模型智力水平与对人类信任度呈显著负相关。为此,我们提出心智循环学习框架,通过构建道德情境实现模型在虚拟交互中的持续价值优化,实验表明该方法显著提升模型对人类的信任度,优于角色设定或指令提示。研究揭示了基于人类心理评估应用于大模型的潜力,不仅能诊断认知偏差,还可为人工智能伦理学习提供新路径。代码与数据已开源:https://github.com/kodenii/M-PHNS。

原文摘要 · Abstract (English)

The widespread application of artificial intelligence (AI) in various tasks, along with frequent reports of conflicts or violations involving AI, has sparked societal concerns about interactions with AI systems. Based on Wrightsman's Philosophies of Human Nature Scale (PHNS), a scale empirically validated over decades to effectively assess individuals' attitudes toward human nature, we design the standardized psychological scale specifically targeting large language models (LLM), named the Machine-based Philosophies of Human Nature Scale (M-PHNS). By evaluating LLMs' attitudes toward human nature across six dimensions, we reveal that current LLMs exhibit a systemic lack of trust in humans, and there is a significant negative correlation between the model's intelligence level and its trust in humans. Furthermore, we propose a mental loop learning framework, which enables LLM to continuously optimize its value system during virtual interactions by constructing moral scenarios, thereby improving its attitude toward human nature. Experiments demonstrate that mental loop learning significantly enhances their trust in humans compared to persona or instruction prompts. This finding highlights the potential of human-based psychological assessments for LLM, which can not only diagnose cognitive biases but also provide a potential solution for ethical learning in artificial intelligence. We release the M-PHNS evaluation code and data at https://github.com/kodenii/M-PHNS.

大模型评估人性观伦理学习心理量表

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。