arXiv:2607.08285cs.AI2026-07

AI评估缺了心理能力维度,影响人机互动效果。

Psychological Competence as a Missing Dimension in AI Evaluation

  • 提出心理胜任力概念,评估AI如何支持用户认知与决策。
  • 强调对话框架、语气、不确定性处理等对用户影响的关键作用。
  • 适合关注AI真实社会影响的研究者与监管者参考。

当前的AI评估体系主要关注准确性、鲁棒性、推理能力和政策合规性等技术表现,这些指标虽重要,但不足以涵盖直接与用户交互的AI系统。随着自然语言交互的AI越来越多地担任顾问、教练、导师和伙伴角色,其回应会影响用户思考方式、情绪理解、信念形成、信任校准和决策过程。因此,评估单位应从模型本身扩展到人机互动整体。本文提出心理胜任力作为缺失的评估维度,定义为面向用户的AI系统在特定情境与目的下,以恰当方式支持用户认知、情绪解读与行为决策的能力。这包括对话中的框架设计、语气、权威感、响应性、不确定性处理及引导策略等属性。现有评估方法虽涉及部分要素,但极少直接衡量这些心理效应。基于行为科学与人机交互研究,本文构建了心理胜任力的概念框架及其核心领域,并建议通过情景化探针、结构化人类评估和模型辅助方法进行测评。我们主张,心理胜任力应成为模型提供方、部署机构、研究人员与监管者共同关注的核心议题。

原文摘要 · Abstract (English)

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

AI评估人机交互心理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。